System-Level Performance Modeling of Photonic In-Memory Computing
Fuente:
arXiv
Saved in:
| Main Authors: | Arockiaraj, Jebacyril, Wijeratne, Sasindu, Sunder, Sugeet, Kaiser, Md Abdullah-Al, Jaiswal, Akhilesh, Jacob, Ajey P., Prasanna, Viktor |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predictive Performance of Photonic SRAM-based In-Memory Computing for Tensor Decomposition
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2024)
by: Wijeratne, Sasindu, et al.
Published: (2024)
X-pSRAM: A Photonic SRAM with Embedded XOR Logic for Ultra-Fast In-Memory Computing
by: Kaiser, Md Abdullah-Al, et al.
Published: (2025)
by: Kaiser, Md Abdullah-Al, et al.
Published: (2025)
Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
by: Zhang, Zuoning, et al.
Published: (2024)
by: Zhang, Zuoning, et al.
Published: (2024)
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
by: Ramachandran, Anvitha, et al.
Published: (2025)
by: Ramachandran, Anvitha, et al.
Published: (2025)
A Unified CPU-GPU Protocol for GNN Training
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
by: Gupta, Neelesh, et al.
Published: (2025)
by: Gupta, Neelesh, et al.
Published: (2025)
A Mixed-Signal Photonic SRAM-based High-Speed Energy-Efficient Photonic Tensor Core with Novel Electro-Optic ADC
by: Kaiser, Md Abdullah-Al, et al.
Published: (2025)
by: Kaiser, Md Abdullah-Al, et al.
Published: (2025)
RIMMS: Runtime Integrated Memory Management System for Heterogeneous Computing
by: Gener, Serhan, et al.
Published: (2025)
by: Gener, Serhan, et al.
Published: (2025)
A New Execution Model and Executor for Adaptively Optimizing the Performance of Parallel Algorithms Using HPX Runtime System
by: Mohammadiporshokooh, Karame, et al.
Published: (2025)
by: Mohammadiporshokooh, Karame, et al.
Published: (2025)
Exploring the Viability of Unikernels for ARM-powered Edge Computing
by: Kaiser, Shahidullah, et al.
Published: (2024)
by: Kaiser, Shahidullah, et al.
Published: (2024)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Performance Trade-offs of High Order Meshless Approximation on Distributed Memory Systems
by: Vehovar, Jon, et al.
Published: (2025)
by: Vehovar, Jon, et al.
Published: (2025)
Exploiting Stragglers in Distributed Computing Systems with Task Grouping
by: Adikari, Tharindu, et al.
Published: (2024)
by: Adikari, Tharindu, et al.
Published: (2024)
Advancing Anomaly Detection in Computational Workflows with Active Learning
by: Raghavan, Krishnan, et al.
Published: (2024)
by: Raghavan, Krishnan, et al.
Published: (2024)
LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks
by: Nader, Noujoud, et al.
Published: (2025)
by: Nader, Noujoud, et al.
Published: (2025)
Design of Energy-Efficient Cross-coupled Differential Photonic-SRAM (pSRAM) Bitcell for High-Speed On-Chip Photonic Memory and Compute Systems
by: Kaiser, Md Abdullah-Al, et al.
Published: (2025)
by: Kaiser, Md Abdullah-Al, et al.
Published: (2025)
Fixing Non-blocking Data Structures for Better Compatibility with Memory Reclamation Schemes
by: Arovi, Md Amit Hasan, et al.
Published: (2025)
by: Arovi, Md Amit Hasan, et al.
Published: (2025)
Flow-Bench: A Dataset for Computational Workflow Anomaly Detection
by: Papadimitriou, George, et al.
Published: (2023)
by: Papadimitriou, George, et al.
Published: (2023)
Understanding the Communication Needs of Asynchronous Many-Task Systems -- A Case Study of HPX+LCI
by: Yan, Jiakun, et al.
Published: (2025)
by: Yan, Jiakun, et al.
Published: (2025)
Edge System Design Using Containers and Unikernels for IoT Applications
by: Kaiser, Shahidullah, et al.
Published: (2024)
by: Kaiser, Shahidullah, et al.
Published: (2024)
Performance of Distributed File Systems on Cloud Computing Environment: An Evaluation for Small-File Problem
by: Duong, Thanh, et al.
Published: (2023)
by: Duong, Thanh, et al.
Published: (2023)
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
by: Zheng, Haoyu, et al.
Published: (2025)
by: Zheng, Haoyu, et al.
Published: (2025)
Overcoming Latency-bound Limitations of Distributed Graph Algorithms using the HPX Runtime System
by: Mohammadiporshokooh, Karame, et al.
Published: (2026)
by: Mohammadiporshokooh, Karame, et al.
Published: (2026)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
by: Lu, Zhengxian, et al.
Published: (2024)
by: Lu, Zhengxian, et al.
Published: (2024)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
by: Wahlgren, Jacob, et al.
Published: (2024)
by: Wahlgren, Jacob, et al.
Published: (2024)
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
by: Miksits, Samuel, et al.
Published: (2024)
by: Miksits, Samuel, et al.
Published: (2024)
Message-Oriented Middleware Systems: Technology Overview
by: Al-Manasrah, Wael, et al.
Published: (2026)
by: Al-Manasrah, Wael, et al.
Published: (2026)
End-to-End and Phase-Level Performance Optimization for Hyperledger Fabric
by: Sollu, Pavan, et al.
Published: (2026)
by: Sollu, Pavan, et al.
Published: (2026)
Computation-Bandwidth-Memory Trade-offs: A Unified Paradigm for AI Infrastructure
by: Fan, Yuankai, et al.
Published: (2025)
by: Fan, Yuankai, et al.
Published: (2025)
Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing
by: Ferdous, S M, et al.
Published: (2024)
by: Ferdous, S M, et al.
Published: (2024)
Shared Virtual Memory: Its Design and Performance Implications for Diverse Applications
by: Cooper, Bennett, et al.
Published: (2024)
by: Cooper, Bennett, et al.
Published: (2024)
Accurate Performance Predictors for Edge Computing Applications
by: Giannakopoulos, Panagiotis, et al.
Published: (2025)
by: Giannakopoulos, Panagiotis, et al.
Published: (2025)
Performance Characterization of Distributed Deep Learning Strategies: A Quantitative Evaluation of DDP, FSDP, and Parameter Server Architectures on GPU Clusters
by: Ovi, Md Sultanul Islam
Published: (2025)
by: Ovi, Md Sultanul Islam
Published: (2025)
Modular Architecture for High-Performance and Low Overhead Data Transfers
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
Similar Items
-
Predictive Performance of Photonic SRAM-based In-Memory Computing for Tensor Decomposition
by: Wijeratne, Sasindu, et al.
Published: (2025) -
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2025) -
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
by: Wijeratne, Sasindu, et al.
Published: (2025) -
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2024) -
X-pSRAM: A Photonic SRAM with Embedded XOR Logic for Ultra-Fast In-Memory Computing
by: Kaiser, Md Abdullah-Al, et al.
Published: (2025)