PCCL: Photonic circuit-switched collective communication for distributed ML
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Abhishek Vijaya, Devraj, Arjun, Singh, Rachee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient AllReduce with Stragglers
by: Devraj, Arjun, et al.
Published: (2025)
by: Devraj, Arjun, et al.
Published: (2025)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
by: Ding, Eric, et al.
Published: (2026)
by: Ding, Eric, et al.
Published: (2026)
Eliminating Hidden Serialization in Multi-Node Megakernel Communication
by: Oh, Byungsoo, et al.
Published: (2026)
by: Oh, Byungsoo, et al.
Published: (2026)
Short-circuiting Rings for Low-Latency AllReduce
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
by: Svedas, Jonas, et al.
Published: (2026)
by: Svedas, Jonas, et al.
Published: (2026)
FlashMoE: Fast Distributed MoE in a Single Kernel
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025)
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025)
A monitoring system for collecting and aggregating metrics from distributed clouds
by: Ranković, Tamara, et al.
Published: (2026)
by: Ranković, Tamara, et al.
Published: (2026)
HPAC-ML: A Programming Model for Embedding ML Surrogates in Scientific Applications
by: Fink, Zane, et al.
Published: (2024)
by: Fink, Zane, et al.
Published: (2024)
Simulating LLM training workloads for heterogeneous compute and network infrastructure
by: Kumar, Sumit, et al.
Published: (2025)
by: Kumar, Sumit, et al.
Published: (2025)
Fog enabled distributed training architecture for federated learning
by: Kumar, Aditya, et al.
Published: (2024)
by: Kumar, Aditya, et al.
Published: (2024)
Declarative Data Pipeline for Large Scale ML Services
by: Yang, Yunzhao, et al.
Published: (2025)
by: Yang, Yunzhao, et al.
Published: (2025)
A Self-Healing and Fault-Tolerant Cloud-based Digital Twin Processing Management Model
by: Saxena, Deepika, et al.
Published: (2025)
by: Saxena, Deepika, et al.
Published: (2025)
Multi-Factor Trust-Driven Secure Communication Model for Cloud-Based Digital Twins
by: Saxena, Deepika, et al.
Published: (2026)
by: Saxena, Deepika, et al.
Published: (2026)
ML-based Adaptive Prefetching and Data Placement for US HEP Systems
by: Karanam, Venkat Sai Suman Lamba, et al.
Published: (2025)
by: Karanam, Venkat Sai Suman Lamba, et al.
Published: (2025)
Gathering of asynchronous robots on circle with limited visibility using finite communication
by: Sharma, Avisek, et al.
Published: (2025)
by: Sharma, Avisek, et al.
Published: (2025)
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
by: Svedas, Jonas, et al.
Published: (2025)
by: Svedas, Jonas, et al.
Published: (2025)
A Performance Analyzer for a Public Cloud's ML-Augmented VM Allocator
by: Bostandoost, Roozbeh, et al.
Published: (2025)
by: Bostandoost, Roozbeh, et al.
Published: (2025)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
by: Mehboob, Talha, et al.
Published: (2025)
by: Mehboob, Talha, et al.
Published: (2025)
NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
by: Yoo, Jinsun, et al.
Published: (2026)
by: Yoo, Jinsun, et al.
Published: (2026)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)
by: Ahmad, Sohaib, et al.
Published: (2024)
ML-ECS: A Collaborative Multimodal Learning Framework for Edge-Cloud Synergies
by: Liu, Yuze, et al.
Published: (2026)
by: Liu, Yuze, et al.
Published: (2026)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
by: Jain, Rutwik, et al.
Published: (2024)
by: Jain, Rutwik, et al.
Published: (2024)
ML-based Modeling to Predict I/O Performance on Different Storage Sub-systems
by: Xu, Yiheng, et al.
Published: (2023)
by: Xu, Yiheng, et al.
Published: (2023)
System-Level Performance Modeling of Photonic In-Memory Computing
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
by: Wolfrath, Joel, et al.
Published: (2025)
by: Wolfrath, Joel, et al.
Published: (2025)
BlazingAML: High-Throughput Anti-Money Laundering (AML) via Multi-Stage Graph Mining
by: Ye, Haojie, et al.
Published: (2026)
by: Ye, Haojie, et al.
Published: (2026)
HeLoCo: Efficient asynchronous low-communication training under data and device heterogeneity
by: Asif, Abdullah Al, et al.
Published: (2026)
by: Asif, Abdullah Al, et al.
Published: (2026)
LLM-assisted Agentic Edge Intelligence Framework
by: Dehury, Chinmaya Kumar, et al.
Published: (2026)
by: Dehury, Chinmaya Kumar, et al.
Published: (2026)
OCEP: An Ontology-Based Complex Event Processing Framework for Healthcare Decision Support in Big Data Analytics
by: Chandra, Ritesh, et al.
Published: (2025)
by: Chandra, Ritesh, et al.
Published: (2025)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
by: Jeon, Beomyeol, et al.
Published: (2024)
by: Jeon, Beomyeol, et al.
Published: (2024)
Predictive Performance of Photonic SRAM-based In-Memory Computing for Tensor Decomposition
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Why Atomicity Matters to AI/ML Infrastructure: Snapshots, Firmware Updates, and the Cost of the Forward-In-Time-Only Category Mistake
by: Borrill, Paul
Published: (2026)
by: Borrill, Paul
Published: (2026)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
by: Tang, Yupeng, et al.
Published: (2023)
by: Tang, Yupeng, et al.
Published: (2023)
Binary integer programming for optimizing ebit cost in distributed quantum circuits with fixed module allocation
by: Cha, Hyunho, et al.
Published: (2025)
by: Cha, Hyunho, et al.
Published: (2025)
Configuration management in the distributed cloud
by: Ranković, Tamara, et al.
Published: (2024)
by: Ranković, Tamara, et al.
Published: (2024)
Snowpark: Performant, Secure, User-Friendly Data Engineering and AI/ML Next To Your Data
by: Baker, Brandon, et al.
Published: (2025)
by: Baker, Brandon, et al.
Published: (2025)
Cold Start Latency in Serverless Computing: A Systematic Review, Taxonomy, and Future Directions
by: Golec, Muhammed, et al.
Published: (2023)
by: Golec, Muhammed, et al.
Published: (2023)
Autonomic Cloud Computing: Research Perspective
by: Gill, Sukhpal Singh
Published: (2015)
by: Gill, Sukhpal Singh
Published: (2015)
Similar Items
-
Efficient AllReduce with Stragglers
by: Devraj, Arjun, et al.
Published: (2025) -
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
by: Kumar, Abhishek Vijaya, et al.
Published: (2024) -
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
by: Ding, Eric, et al.
Published: (2026) -
Eliminating Hidden Serialization in Multi-Node Megakernel Communication
by: Oh, Byungsoo, et al.
Published: (2026) -
Short-circuiting Rings for Low-Latency AllReduce
by: Hammer, Sarah-Michelle, et al.
Published: (2025)