SAKURAONE: An Open Ethernet-Based AI HPC System and Its Observed Workload Dynamics in a Single-Tenant LLM Development Environment
Fuente:
arXiv
Saved in:
| Main Authors: | Konishi, Fumikazu, Tsubouchi, Yuuki, Tsuruta, Hirofumi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAKURAONE: Empowering Transparent and Open AI Platforms through Private-Sector HPC Investment in Japan
by: Konishi, Fumikazu
Published: (2025)
by: Konishi, Fumikazu
Published: (2025)
Closing the HPC-Cloud Convergence Gap: Multi-Tenant Slingshot RDMA for Kubernetes
by: Friese, Philipp A., et al.
Published: (2025)
by: Friese, Philipp A., et al.
Published: (2025)
Causal Inference for Quantifying Noisy Neighbor Effects in Multi-Tenant Cloud Environments
by: Schiavo, Philipe S., et al.
Published: (2026)
by: Schiavo, Philipe S., et al.
Published: (2026)
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
by: Zambianco, Marco, et al.
Published: (2025)
by: Zambianco, Marco, et al.
Published: (2025)
Palladium: A DPU-enabled Multi-Tenant Serverless Cloud over Zero-copy Multi-node RDMA Fabrics
by: Qi, Shixiong, et al.
Published: (2025)
by: Qi, Shixiong, et al.
Published: (2025)
Study of Workload Interference with Intelligent Routing on Dragonfly
by: Kang, Yao, et al.
Published: (2024)
by: Kang, Yao, et al.
Published: (2024)
To Stream or Not to Stream: Towards A Quantitative Model for Remote HPC Processing Decisions
by: Castro, Flavio, et al.
Published: (2025)
by: Castro, Flavio, et al.
Published: (2025)
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
by: Warraich, Ertza, et al.
Published: (2025)
by: Warraich, Ertza, et al.
Published: (2025)
The Bilateral Efficiency of Ethernet: Recalibrating Metcalfe and Boggs After Fifty Years
by: Borrill, Paul
Published: (2026)
by: Borrill, Paul
Published: (2026)
Ultra Ethernet's Design Principles and Architectural Innovations
by: Hoefler, Torsten, et al.
Published: (2025)
by: Hoefler, Torsten, et al.
Published: (2025)
Dynamic Edge Server Selection in Time-Varying Environments: A Reliability-Aware Predictive Approach
by: Burbano, Jaime Sebastian, et al.
Published: (2025)
by: Burbano, Jaime Sebastian, et al.
Published: (2025)
A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks
by: Han, Mingqi, et al.
Published: (2026)
by: Han, Mingqi, et al.
Published: (2026)
An Open-Source Experimentation Framework for the Edge Cloud Continuum
by: Koukis, Georgios, et al.
Published: (2024)
by: Koukis, Georgios, et al.
Published: (2024)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
by: Yang, Zheming, et al.
Published: (2024)
by: Yang, Zheming, et al.
Published: (2024)
An Open-Source Fast Parallel Routing Approach for Commercial FPGAs
by: Zang, Xinshi, et al.
Published: (2024)
by: Zang, Xinshi, et al.
Published: (2024)
The Fog Development Kit: A Development Platform for SDN-based Edge-Fog Systems
by: Powell, Colton, et al.
Published: (2019)
by: Powell, Colton, et al.
Published: (2019)
Multi-stage Flow Scheduling for LLM Serving
by: Sun, Yijun, et al.
Published: (2026)
by: Sun, Yijun, et al.
Published: (2026)
Recursive Offloading for LLM Serving in Multi-tier Networks
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
A Density-Delay Law for Stable Event-Driven State Progression in Open Distributed Systems
by: Chen, Bin, et al.
Published: (2026)
by: Chen, Bin, et al.
Published: (2026)
COREC: Concurrent Non-Blocking Single-Queue Receive Driver for Low Latency Networking
by: Faltelli, Marco, et al.
Published: (2024)
by: Faltelli, Marco, et al.
Published: (2024)
Contextual Chain: Single-State Ledger Design for Mobile/IoT Networks with Frequent Partitions
by: Kim, Song-Ju
Published: (2026)
by: Kim, Song-Ju
Published: (2026)
Collective Communication Profiling of Modern-day Machine Learning Workloads
by: Gupta, Jit, et al.
Published: (2025)
by: Gupta, Jit, et al.
Published: (2025)
A Multi-Cloud Framework for Zero-Trust Workload Authentication
by: Deochake, Saurabh, et al.
Published: (2025)
by: Deochake, Saurabh, et al.
Published: (2025)
Solving AI Foundational Model Latency with Telco Infrastructure
by: Barros, Sebastian
Published: (2025)
by: Barros, Sebastian
Published: (2025)
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
by: Ding, Eric, et al.
Published: (2026)
by: Ding, Eric, et al.
Published: (2026)
NET4EXA: Pioneering the Future of Interconnects for Supercomputing and AI
by: Martinelli, Michele, et al.
Published: (2026)
by: Martinelli, Michele, et al.
Published: (2026)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025)
by: Du, Chengze, et al.
Published: (2025)
MOFCO: Mobility- and Migration-Aware Task Offloading in Three-Layer Fog Computing Environments
by: Mahdizadeh, Soheil, et al.
Published: (2025)
by: Mahdizadeh, Soheil, et al.
Published: (2025)
Semantic-Aware LLM Orchestration for Proactive Resource Management in Predictive Digital Twin Vehicular Networks
by: Ahmadpanah, Seyed Hossein
Published: (2025)
by: Ahmadpanah, Seyed Hossein
Published: (2025)
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
by: Luo, Haoxiang, et al.
Published: (2025)
by: Luo, Haoxiang, et al.
Published: (2025)
CRAFT: Latency and Cost-Aware Genetic-Based Framework for Node Placement in Edge-Fog Environments
by: Mahdizadeh, Soheil, et al.
Published: (2025)
by: Mahdizadeh, Soheil, et al.
Published: (2025)
OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
by: Tan, Xin, et al.
Published: (2026)
by: Tan, Xin, et al.
Published: (2026)
GORGO: Maximizing KV-Cache Reuse While Minimizing Network Latency in Cross-Region LLM Load Balancing
by: Toniolo, Alessio Ricci, et al.
Published: (2026)
by: Toniolo, Alessio Ricci, et al.
Published: (2026)
FlowTracer: A Tool for Uncovering Network Path Usage Imbalance in AI Training Clusters
by: Jamil, Hasibul, et al.
Published: (2024)
by: Jamil, Hasibul, et al.
Published: (2024)
Hurry: Dynamic Collaborative Framework For Low-orbit Mega-Constellation Data Downloading
by: Luo, Handong, et al.
Published: (2024)
by: Luo, Handong, et al.
Published: (2024)
Dynamic DAG-Application Scheduling for Multi-Tier Edge Computing in Heterogeneous Networks
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Dynamic Hierarchical Birkhoff-von Neumann Decomposition for All-to-All GPU Communication
by: Wu, Yen-Chieh, et al.
Published: (2026)
by: Wu, Yen-Chieh, et al.
Published: (2026)
DRDST: Low-latency DAG Consensus through Robust Dynamic Sharding and Tree-broadcasting for IoV
by: Chen, Runhua, et al.
Published: (2024)
by: Chen, Runhua, et al.
Published: (2024)
Accelerating Stable Matching between Workers and Spatial-Temporal Tasks for Dynamic MCS: A Stagewise Service Trading Approach
by: Qi, Houyi, et al.
Published: (2025)
by: Qi, Houyi, et al.
Published: (2025)
Similar Items
-
SAKURAONE: Empowering Transparent and Open AI Platforms through Private-Sector HPC Investment in Japan
by: Konishi, Fumikazu
Published: (2025) -
Closing the HPC-Cloud Convergence Gap: Multi-Tenant Slingshot RDMA for Kubernetes
by: Friese, Philipp A., et al.
Published: (2025) -
Causal Inference for Quantifying Noisy Neighbor Effects in Multi-Tenant Cloud Environments
by: Schiavo, Philipe S., et al.
Published: (2026) -
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
by: Zambianco, Marco, et al.
Published: (2025) -
Palladium: A DPU-enabled Multi-Tenant Serverless Cloud over Zero-copy Multi-node RDMA Fabrics
by: Qi, Shixiong, et al.
Published: (2025)