Toward Self-Healing Networks-on-Chip: RL-Driven Routing in 2D Torus Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | Charrwi, Mohammad Walid, Hussain, Zaid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Resilient Packet Forwarding: A Reinforcement Learning Approach to Routing in Gaussian Interconnected Networks with Clustered Faults
by: Charrwi, Mohammad Walid, et al.
Published: (2025)
by: Charrwi, Mohammad Walid, et al.
Published: (2025)
Deep Reinforcement Learning for Fault-Adaptive Routing in Eisenstein-Jacobi Interconnection Topologies
by: Charrwi, Mohammad Walid, et al.
Published: (2026)
by: Charrwi, Mohammad Walid, et al.
Published: (2026)
Resilient Auto-Scaling of Microservice Architectures with Efficient Resource Management
by: Ahmad, Hussain, et al.
Published: (2025)
by: Ahmad, Hussain, et al.
Published: (2025)
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
by: Tang, Ding, et al.
Published: (2025)
by: Tang, Ding, et al.
Published: (2025)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
by: Yu, Shibo, et al.
Published: (2025)
by: Yu, Shibo, et al.
Published: (2025)
Quantifying the Carbon Reduction of DAG Workloads: A Job Shop Scheduling Perspective
by: Bostandoost, Roozbeh, et al.
Published: (2025)
by: Bostandoost, Roozbeh, et al.
Published: (2025)
A Self-Healing and Fault-Tolerant Cloud-based Digital Twin Processing Management Model
by: Saxena, Deepika, et al.
Published: (2025)
by: Saxena, Deepika, et al.
Published: (2025)
Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving
by: Yoshimura, Takeshi, et al.
Published: (2026)
by: Yoshimura, Takeshi, et al.
Published: (2026)
RL over Commodity Networks: Overcoming the Bandwidth Barrier with Lossless Sparse Deltas
by: Ruan, Chaoyi, et al.
Published: (2026)
by: Ruan, Chaoyi, et al.
Published: (2026)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
by: Liu, Yiqi, et al.
Published: (2026)
by: Liu, Yiqi, et al.
Published: (2026)
Towards Learning-Augmented Peer-to-Peer Networks: Self-Stabilizing Graph Linearization with Untrusted Advice
by: Aradhya, Vijeth, et al.
Published: (2025)
by: Aradhya, Vijeth, et al.
Published: (2025)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
by: Li, Zixuan, et al.
Published: (2026)
by: Li, Zixuan, et al.
Published: (2026)
SHARE: Optimizing Secure Hub Allocation and Routing Efficiency in Payment Channel Networks
by: Yang, Lingxiao, et al.
Published: (2025)
by: Yang, Lingxiao, et al.
Published: (2025)
Achieving High-Performance Fault-Tolerant Routing in HyperX Interconnection Networks
by: Camarero, Cristóbal, et al.
Published: (2024)
by: Camarero, Cristóbal, et al.
Published: (2024)
Octopus: Experiences with a Hybrid Event-Driven Architecture for Distributed Scientific Computing
by: Pan, Haochen, et al.
Published: (2024)
by: Pan, Haochen, et al.
Published: (2024)
SynergAI: Edge-to-Cloud Synergy for Architecture-Driven High-Performance Orchestration for AI Inference
by: Stathopoulou, Foteini, et al.
Published: (2025)
by: Stathopoulou, Foteini, et al.
Published: (2025)
Data Augmentation and Convolutional Network Architecture Influence on Distributed Learning
by: Jansen, Victor Forattini, et al.
Published: (2026)
by: Jansen, Victor Forattini, et al.
Published: (2026)
Polar: Agentic RL on Any Harness at Scale
by: Xu, Binfeng, et al.
Published: (2026)
by: Xu, Binfeng, et al.
Published: (2026)
Adventures with Grace Hopper AI Super Chip and the National Research Platform
by: Hurt, J. Alex, et al.
Published: (2024)
by: Hurt, J. Alex, et al.
Published: (2024)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
by: Shi, Long, et al.
Published: (2025)
by: Shi, Long, et al.
Published: (2025)
CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters
by: Hanafy, Walid A., et al.
Published: (2025)
by: Hanafy, Walid A., et al.
Published: (2025)
Adaptive Self-Organization in Anonymous Dynamic Networks
by: Parzych, Garrett, et al.
Published: (2026)
by: Parzych, Garrett, et al.
Published: (2026)
HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
by: He, Yongjun, et al.
Published: (2025)
by: He, Yongjun, et al.
Published: (2025)
Schedule-Level Shared-Prefix Reuse for LLM RL Training
by: Li, Pengbo, et al.
Published: (2026)
by: Li, Pengbo, et al.
Published: (2026)
Polylog-Competitive Deterministic Local Routing and Scheduling
by: Haeupler, Bernhard, et al.
Published: (2024)
by: Haeupler, Bernhard, et al.
Published: (2024)
LLM-Driven Intent-Based Privacy-Aware Orchestration Across the Cloud-Edge Continuum
by: Su, Zijie, et al.
Published: (2026)
by: Su, Zijie, et al.
Published: (2026)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
Unicron: Economizing Self-Healing LLM Training at Scale
by: He, Tao, et al.
Published: (2023)
by: He, Tao, et al.
Published: (2023)
MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
by: Tang, Xinru, et al.
Published: (2025)
by: Tang, Xinru, et al.
Published: (2025)
BeRGeR: Byzantine-Robust Geometric Routing
by: Zaz, Brown, et al.
Published: (2024)
by: Zaz, Brown, et al.
Published: (2024)
KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Unleashing Efficient Asynchronous RL Post-Training via Staleness-Constrained Rollout Coordination
by: Li, Haoyang, et al.
Published: (2026)
by: Li, Haoyang, et al.
Published: (2026)
HiRL: Hierarchical Reinforcement Learning for Coordinated Resource Management in Heterogeneous Edge Computing
by: Zhu, Jianyong, et al.
Published: (2026)
by: Zhu, Jianyong, et al.
Published: (2026)
Universal Finite-State and Self-Stabilizing Computation in Anonymous Dynamic Networks
by: Di Luna, Giuseppe A., et al.
Published: (2024)
by: Di Luna, Giuseppe A., et al.
Published: (2024)
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
by: Dai, Yinwei, et al.
Published: (2025)
by: Dai, Yinwei, et al.
Published: (2025)
Resource Allocation of Industry 4.0 Micro-Service Applications across Serverless Fog Federation
by: Hussain, Razin Farhan, et al.
Published: (2024)
by: Hussain, Razin Farhan, et al.
Published: (2024)
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
by: Wang, Zhixin, et al.
Published: (2025)
by: Wang, Zhixin, et al.
Published: (2025)
Vanishing Variance Problem in Fully Decentralized Neural-Network Systems
by: Tian, Yongding, et al.
Published: (2024)
by: Tian, Yongding, et al.
Published: (2024)
Communication Lower Bounds and Algorithms for Sketching with Random Dense Matrices
by: Daas, Hussam Al, et al.
Published: (2026)
by: Daas, Hussam Al, et al.
Published: (2026)
Similar Items
-
Resilient Packet Forwarding: A Reinforcement Learning Approach to Routing in Gaussian Interconnected Networks with Clustered Faults
by: Charrwi, Mohammad Walid, et al.
Published: (2025) -
Deep Reinforcement Learning for Fault-Adaptive Routing in Eisenstein-Jacobi Interconnection Topologies
by: Charrwi, Mohammad Walid, et al.
Published: (2026) -
Resilient Auto-Scaling of Microservice Architectures with Efficient Resource Management
by: Ahmad, Hussain, et al.
Published: (2025) -
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
by: Tang, Ding, et al.
Published: (2025) -
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
by: Yu, Shibo, et al.
Published: (2025)