Next-generation Probabilistic Computing Hardware with 3D MOSAICs, Illusion Scale-up, and Co-design
Fuente:
arXiv
Guardado en:
| Autores principales: | Srimani, Tathagata, Radway, Robert, Mohseni, Masoud, Çamsarı, Kerem, Mitra, Subhasish |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
por: Qu, Huanyu, et al.
Publicado: (2025)
por: Qu, Huanyu, et al.
Publicado: (2025)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
por: Feng, Weigang, et al.
Publicado: (2025)
por: Feng, Weigang, et al.
Publicado: (2025)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
por: Kim, Hyeseong, et al.
Publicado: (2026)
por: Kim, Hyeseong, et al.
Publicado: (2026)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
por: Kwak, Hyunseok, et al.
Publicado: (2025)
por: Kwak, Hyunseok, et al.
Publicado: (2025)
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
por: Garg, Raveesh, et al.
Publicado: (2023)
por: Garg, Raveesh, et al.
Publicado: (2023)
Tascade: Hardware Support for Atomic-free, Asynchronous and Efficient Reduction Trees
por: Orenes-Vera, Marcelo, et al.
Publicado: (2023)
por: Orenes-Vera, Marcelo, et al.
Publicado: (2023)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
por: Adnan, Muhammad, et al.
Publicado: (2024)
por: Adnan, Muhammad, et al.
Publicado: (2024)
Switchboard: An Open-Source Framework for Modular Simulation of Large Hardware Systems
por: Herbst, Steven, et al.
Publicado: (2024)
por: Herbst, Steven, et al.
Publicado: (2024)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
por: Li, Jiamin, et al.
Publicado: (2025)
por: Li, Jiamin, et al.
Publicado: (2025)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
por: Qiu, Tong Dong, et al.
Publicado: (2023)
por: Qiu, Tong Dong, et al.
Publicado: (2023)
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
por: Wang, Jing, et al.
Publicado: (2025)
por: Wang, Jing, et al.
Publicado: (2025)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
por: Zhang, Yichao, et al.
Publicado: (2026)
por: Zhang, Yichao, et al.
Publicado: (2026)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
por: Zhou, Zhuoshan, et al.
Publicado: (2026)
por: Zhou, Zhuoshan, et al.
Publicado: (2026)
Memory-Centric Computing: Solving Computing's Memory Problem
por: Mutlu, Onur, et al.
Publicado: (2025)
por: Mutlu, Onur, et al.
Publicado: (2025)
NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data Processing
por: Zou, Cheng, et al.
Publicado: (2026)
por: Zou, Cheng, et al.
Publicado: (2026)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
por: Zhang, Chen, et al.
Publicado: (2026)
por: Zhang, Chen, et al.
Publicado: (2026)
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
por: Bai, Zhenyu, et al.
Publicado: (2025)
por: Bai, Zhenyu, et al.
Publicado: (2025)
Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
por: Meyer, Marius, et al.
Publicado: (2024)
por: Meyer, Marius, et al.
Publicado: (2024)
Context-aware Simopt-Power: Using structural data with simulation metadata to optimise FPGA designs
por: Wadhwa, Eashan, et al.
Publicado: (2026)
por: Wadhwa, Eashan, et al.
Publicado: (2026)
Revisiting Computational Storage for Data Integrity and Security
por: Shi, Chao, et al.
Publicado: (2025)
por: Shi, Chao, et al.
Publicado: (2025)
TeraPool-SDR: An 1.89TOPS 1024 RV-Cores 4MiB Shared-L1 Cluster for Next-Generation Open-Source Software-Defined Radios
por: Zhang, Yichao, et al.
Publicado: (2024)
por: Zhang, Yichao, et al.
Publicado: (2024)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
por: Sirjani, Mohammad Sadegh, et al.
Publicado: (2025)
por: Sirjani, Mohammad Sadegh, et al.
Publicado: (2025)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
por: Mutlu, Onur, et al.
Publicado: (2024)
por: Mutlu, Onur, et al.
Publicado: (2024)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
por: Punniyamurthy, Kishore, et al.
Publicado: (2023)
por: Punniyamurthy, Kishore, et al.
Publicado: (2023)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
por: Agrawal, Anirudha, et al.
Publicado: (2024)
por: Agrawal, Anirudha, et al.
Publicado: (2024)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
por: Zhang, Qijun, et al.
Publicado: (2026)
por: Zhang, Qijun, et al.
Publicado: (2026)
Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods
por: Fatima, Amel, et al.
Publicado: (2026)
por: Fatima, Amel, et al.
Publicado: (2026)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
por: Ma, Ke, et al.
Publicado: (2025)
por: Ma, Ke, et al.
Publicado: (2025)
BlockAMC: Scalable In-Memory Analog Matrix Computing for Solving Linear Systems
por: Pan, Lunshuai, et al.
Publicado: (2024)
por: Pan, Lunshuai, et al.
Publicado: (2024)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
por: Li, Ming, et al.
Publicado: (2024)
por: Li, Ming, et al.
Publicado: (2024)
How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
por: Wang, Yuqing, et al.
Publicado: (2025)
por: Wang, Yuqing, et al.
Publicado: (2025)
UniFormer: Unified and Efficient Transformer for Reasoning Across General and Custom Computing
por: Ran, Zhuoheng, et al.
Publicado: (2025)
por: Ran, Zhuoheng, et al.
Publicado: (2025)
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
por: Eleftherakis, Panagiotis-Eleftherios, et al.
Publicado: (2026)
por: Eleftherakis, Panagiotis-Eleftherios, et al.
Publicado: (2026)
RapidOMS: FPGA-based Open Modification Spectral Library Searching with HD Computing
por: Pinge, Sumukh, et al.
Publicado: (2024)
por: Pinge, Sumukh, et al.
Publicado: (2024)
FpgaHub: Fpga-centric Hyper-heterogeneous Computing Platform for Big Data Analytics
por: Wang, Zeke, et al.
Publicado: (2025)
por: Wang, Zeke, et al.
Publicado: (2025)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
por: Colagrande, Luca, et al.
Publicado: (2026)
por: Colagrande, Luca, et al.
Publicado: (2026)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
por: Yuksel, Ismail Emir, et al.
Publicado: (2023)
por: Yuksel, Ismail Emir, et al.
Publicado: (2023)
Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State Drives
por: Nadig, Rakesh, et al.
Publicado: (2026)
por: Nadig, Rakesh, et al.
Publicado: (2026)
SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence Analysis
por: Ghiasi, Nika Mansouri, et al.
Publicado: (2025)
por: Ghiasi, Nika Mansouri, et al.
Publicado: (2025)
CUCo: An Agentic Framework for Compute and Communication Co-design
por: Hu, Bodun, et al.
Publicado: (2026)
por: Hu, Bodun, et al.
Publicado: (2026)
Ejemplares similares
-
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
por: Qu, Huanyu, et al.
Publicado: (2025) -
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
por: Feng, Weigang, et al.
Publicado: (2025) -
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
por: Kim, Hyeseong, et al.
Publicado: (2026) -
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
por: Kwak, Hyunseok, et al.
Publicado: (2025) -
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
por: Garg, Raveesh, et al.
Publicado: (2023)