Accelerating Frontier MoE Training with 3D Integrated Optics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bernadskiy, Mikhail, Carson, Peter, Graham, Thomas, Groves, Taylor, Lee, Ho John, Yeh, Eric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
von: Li, Yinrong, et al.
Veröffentlicht: (2026)
von: Li, Yinrong, et al.
Veröffentlicht: (2026)
A Comprehensive Survey on Orbital Edge Computing: Systems, Applications, and Algorithms
von: Wu, Changhao, et al.
Veröffentlicht: (2023)
von: Wu, Changhao, et al.
Veröffentlicht: (2023)
Implementation and Evaluation of Fast Raft for Hierarchical Consensus
von: Melnychuk, Anton, et al.
Veröffentlicht: (2025)
von: Melnychuk, Anton, et al.
Veröffentlicht: (2025)
Rendezvous and Merging for Two Metamorphic Robotic Systems without Global Compass
von: Yamada, Ryonosuke, et al.
Veröffentlicht: (2024)
von: Yamada, Ryonosuke, et al.
Veröffentlicht: (2024)
Model-driven development of data intensive applications over cloud resources
von: Tolosana-Calasanz, Rafael, et al.
Veröffentlicht: (2024)
von: Tolosana-Calasanz, Rafael, et al.
Veröffentlicht: (2024)
The $qs$ Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
von: Adhinarayanan, Vignesh, et al.
Veröffentlicht: (2026)
von: Adhinarayanan, Vignesh, et al.
Veröffentlicht: (2026)
Coordinated Reinforcement Learning Prefetching Architecture for Multicore Systems
von: Siddiqui, Mohammed Humaid, et al.
Veröffentlicht: (2025)
von: Siddiqui, Mohammed Humaid, et al.
Veröffentlicht: (2025)
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
von: Sada, Mohammad Firas, et al.
Veröffentlicht: (2025)
von: Sada, Mohammad Firas, et al.
Veröffentlicht: (2025)
Next-Generation Event-Driven Architectures: Performance, Scalability, and Intelligent Orchestration Across Messaging Frameworks
von: Arafat, Jahidul, et al.
Veröffentlicht: (2025)
von: Arafat, Jahidul, et al.
Veröffentlicht: (2025)
SCION: Size-aware Policy Orchestration for Nonstationary Object Caches (Long Paper Version)
von: Wang, Qizhi
Veröffentlicht: (2026)
von: Wang, Qizhi
Veröffentlicht: (2026)
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
von: Goldman, Amos, et al.
Veröffentlicht: (2026)
von: Goldman, Amos, et al.
Veröffentlicht: (2026)
Tetris: An SLA-aware Application Placement Strategy in the Edge-Cloud Continuum
von: Almeida, Lucas, et al.
Veröffentlicht: (2025)
von: Almeida, Lucas, et al.
Veröffentlicht: (2025)
Efficient and Scalable Architecture for Multiple-chip Implementation of Simulated Bifurcation Machines
von: Kashimata, Tomoya, et al.
Veröffentlicht: (2023)
von: Kashimata, Tomoya, et al.
Veröffentlicht: (2023)
Light Cone Consistency: Toward a Unified Theory of Consistency in Message-Passing Systems
von: Landers, Rob, et al.
Veröffentlicht: (2026)
von: Landers, Rob, et al.
Veröffentlicht: (2026)
CRDT-Based Game State Synchronization in Peer-to-Peer VR
von: Dantas, Abel, et al.
Veröffentlicht: (2025)
von: Dantas, Abel, et al.
Veröffentlicht: (2025)
Scalable Engine and the Performance of Different LLM Models in a SLURM based HPC architecture
von: Luiz, Anderson de Lima, et al.
Veröffentlicht: (2025)
von: Luiz, Anderson de Lima, et al.
Veröffentlicht: (2025)
Replication in Graph Partitioning and Scheduling Problems
von: Papp, Pál András, et al.
Veröffentlicht: (2026)
von: Papp, Pál András, et al.
Veröffentlicht: (2026)
GPU-Initiated Networking for NCCL
von: Hamidouche, Khaled, et al.
Veröffentlicht: (2025)
von: Hamidouche, Khaled, et al.
Veröffentlicht: (2025)
Representation Similarity: A Better Guidance of DNN Layer Sharing for Edge Computing without Training
von: Cao, Bryan Bo, et al.
Veröffentlicht: (2024)
von: Cao, Bryan Bo, et al.
Veröffentlicht: (2024)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
von: Li, Zonghang, et al.
Veröffentlicht: (2024)
von: Li, Zonghang, et al.
Veröffentlicht: (2024)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
von: Jung, Myoungsoo
Veröffentlicht: (2025)
von: Jung, Myoungsoo
Veröffentlicht: (2025)
Reproduction Research of FSA-Benchmark
von: Ludolf, Joshua, et al.
Veröffentlicht: (2024)
von: Ludolf, Joshua, et al.
Veröffentlicht: (2024)
Efficient Multi-Processor Scheduling in Increasingly Realistic Models
von: Papp, Pál András, et al.
Veröffentlicht: (2024)
von: Papp, Pál András, et al.
Veröffentlicht: (2024)
Big Data Workload Profiling for Energy-Aware Cloud Resource Management
von: Parikh, Milan, et al.
Veröffentlicht: (2026)
von: Parikh, Milan, et al.
Veröffentlicht: (2026)
SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
von: Li, Xiangchen, et al.
Veröffentlicht: (2025)
von: Li, Xiangchen, et al.
Veröffentlicht: (2025)
A simple protocol to automate the executing, scaling, and reconfiguration of Cloud-Native Apps
von: Ambroszkiewicz, Stanislaw, et al.
Veröffentlicht: (2023)
von: Ambroszkiewicz, Stanislaw, et al.
Veröffentlicht: (2023)
PROV-AGENT: Unified Provenance for Tracking AI Agent Interactions in Agentic Workflows
von: Souza, Renan, et al.
Veröffentlicht: (2025)
von: Souza, Renan, et al.
Veröffentlicht: (2025)
LLM Agents for Interactive Workflow Provenance: Reference Architecture and Evaluation Methodology
von: Souza, Renan, et al.
Veröffentlicht: (2025)
von: Souza, Renan, et al.
Veröffentlicht: (2025)
Quantifying the Performance Gap for Simple Versus Optimal Dynamic Server Allocation Policies
von: Carlsson, Niklas, et al.
Veröffentlicht: (2025)
von: Carlsson, Niklas, et al.
Veröffentlicht: (2025)
FPGA-Accelerated Lock Management and Transaction Processing: Architecture, Optimization, and Design Space Exploration
von: Zhu, Shien, et al.
Veröffentlicht: (2026)
von: Zhu, Shien, et al.
Veröffentlicht: (2026)
Decentralized Task Scheduling in Distributed Systems: A Deep Reinforcement Learning Approach
von: John, Daniel Benniah
Veröffentlicht: (2026)
von: John, Daniel Benniah
Veröffentlicht: (2026)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
SAFELearning: Enable Backdoor Detectability In Federated Learning With Secure Aggregation
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2021)
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2021)
Efficiently Scheduling Parallel DAG Tasks on Identical Multiprocessors
von: Lendve, Shardul, et al.
Veröffentlicht: (2024)
von: Lendve, Shardul, et al.
Veröffentlicht: (2024)
Characterising resource management performance in Kubernetes
von: Medel, Víctor, et al.
Veröffentlicht: (2024)
von: Medel, Víctor, et al.
Veröffentlicht: (2024)
Learning Interpretable Scheduling Algorithms for Data Processing Clusters
von: Hu, Zhibo, et al.
Veröffentlicht: (2024)
von: Hu, Zhibo, et al.
Veröffentlicht: (2024)
Federated Learning in Adversarial Environments: Testbed Design and Poisoning Resilience in Cybersecurity
von: Huang, Hao Jian, et al.
Veröffentlicht: (2024)
von: Huang, Hao Jian, et al.
Veröffentlicht: (2024)
Lincoln AI Computing Survey (LAICS) and Trends
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
Reexamining Paradigms of End-to-End Data Movement
von: Fang, Chin, et al.
Veröffentlicht: (2025)
von: Fang, Chin, et al.
Veröffentlicht: (2025)
FCDP: Fully Cached Data Parallel for Communication-Avoiding Large-Scale Training
von: Park, Gyeongseo, et al.
Veröffentlicht: (2026)
von: Park, Gyeongseo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
von: Li, Yinrong, et al.
Veröffentlicht: (2026) -
A Comprehensive Survey on Orbital Edge Computing: Systems, Applications, and Algorithms
von: Wu, Changhao, et al.
Veröffentlicht: (2023) -
Implementation and Evaluation of Fast Raft for Hierarchical Consensus
von: Melnychuk, Anton, et al.
Veröffentlicht: (2025) -
Rendezvous and Merging for Two Metamorphic Robotic Systems without Global Compass
von: Yamada, Ryonosuke, et al.
Veröffentlicht: (2024) -
Model-driven development of data intensive applications over cloud resources
von: Tolosana-Calasanz, Rafael, et al.
Veröffentlicht: (2024)