BatchWeave: A Consistent Object-Store-Native Data Plane for Large Foundation Model Training
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Ting, Zhang, Junjie, Yan, Xiao, Zhang, Songxin, Song, Zhuoyang, Xi, Jingyi, Mao, Zunyao, Jing, Bingyi, Zhang, Jiaxing, Xie, Zejian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SkyStore: Cost-Optimized Object Storage Across Regions and Clouds
by: Liu, Shu, et al.
Published: (2025)
by: Liu, Shu, et al.
Published: (2025)
PolarStore: High-Performance Data Compression for Large-Scale Cloud-Native Databases
by: Hu, Qingda, et al.
Published: (2025)
by: Hu, Qingda, et al.
Published: (2025)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
by: Shi, Ge, et al.
Published: (2025)
by: Shi, Ge, et al.
Published: (2025)
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
by: Kang, Xueze, et al.
Published: (2025)
by: Kang, Xueze, et al.
Published: (2025)
A Reinforcement Learning Based Backfilling Strategy for HPC Batch Jobs
by: Kolker-Hicks, Elliot, et al.
Published: (2024)
by: Kolker-Hicks, Elliot, et al.
Published: (2024)
GPU-Accelerated Batch-Dynamic Subgraph Matching
by: Qiu, Linshan, et al.
Published: (2024)
by: Qiu, Linshan, et al.
Published: (2024)
An Artificial Intelligence Framework for Joint Structural-Temporal Load Forecasting in Cloud Native Platforms
by: Zhang, Qingyuan
Published: (2026)
by: Zhang, Qingyuan
Published: (2026)
Object Abstraction To Streamline Edge-Cloud-Native Application Development
by: Lertpongrujikorn, Pawissanutt
Published: (2025)
by: Lertpongrujikorn, Pawissanutt
Published: (2025)
UFO3: Weaving the Digital Agent Galaxy
by: Zhang, Chaoyun, et al.
Published: (2025)
by: Zhang, Chaoyun, et al.
Published: (2025)
Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
by: Sakip, Akhmed, et al.
Published: (2026)
by: Sakip, Akhmed, et al.
Published: (2026)
Accelerating Python Applications with Dask and ProxyStore
by: Pauloski, J. Gregory, et al.
Published: (2024)
by: Pauloski, J. Gregory, et al.
Published: (2024)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
by: Chen, Jiabin, et al.
Published: (2024)
by: Chen, Jiabin, et al.
Published: (2024)
CausalMesh: A Formally Verified Causally Consistent Distributed Cache with Support for Client Migration
by: Zhang, Haoran, et al.
Published: (2025)
by: Zhang, Haoran, et al.
Published: (2025)
A Simulated Annealing Approach to Identical Parallel Machine Scheduling
by: Li, Jiaxing, et al.
Published: (2024)
by: Li, Jiaxing, et al.
Published: (2024)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
by: Bai, Fengyao, et al.
Published: (2026)
by: Bai, Fengyao, et al.
Published: (2026)
CDFGNN: a Systematic Design of Cache-based Distributed Full-Batch Graph Neural Network Training with Communication Reduction
by: Zhang, Shuai, et al.
Published: (2024)
by: Zhang, Shuai, et al.
Published: (2024)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
by: Chen, Qiaoling, et al.
Published: (2026)
by: Chen, Qiaoling, et al.
Published: (2026)
WOC: Dual-Path Weighted Object Consensus Made Efficient
by: Fonseca, Tanisha, et al.
Published: (2025)
by: Fonseca, Tanisha, et al.
Published: (2025)
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
by: Hu, Zhisheng, et al.
Published: (2025)
by: Hu, Zhisheng, et al.
Published: (2025)
EACO-RAG: Towards Distributed Tiered LLM Deployment using Edge-Assisted and Collaborative RAG with Adaptive Knowledge Update
by: Li, Jiaxing, et al.
Published: (2024)
by: Li, Jiaxing, et al.
Published: (2024)
User Experiences with MPI RMA and ULFM in a Resilient Key-Value Store Implementation
by: Fohry, Claudia, et al.
Published: (2026)
by: Fohry, Claudia, et al.
Published: (2026)
Federated Learning Using Coupled Tensor Train Decomposition
by: Zhang, Xiangtao, et al.
Published: (2024)
by: Zhang, Xiangtao, et al.
Published: (2024)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
DynoStore: A wide-area distribution system for the management of data over heterogeneous storage
by: Sanchez-Gallegos, Dante D., et al.
Published: (2025)
by: Sanchez-Gallegos, Dante D., et al.
Published: (2025)
Beyond Pre-Training: The Full Lifecycle of Foundation Models on HPC Systems
by: Conciatore, Dino, et al.
Published: (2026)
by: Conciatore, Dino, et al.
Published: (2026)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
by: Gu, Diandian, et al.
Published: (2024)
by: Gu, Diandian, et al.
Published: (2024)
Adaptable TeaStore
by: Bliudze, Simon, et al.
Published: (2024)
by: Bliudze, Simon, et al.
Published: (2024)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
by: Xu, Yaodan, et al.
Published: (2025)
by: Xu, Yaodan, et al.
Published: (2025)
Are Your Epochs Too Epic? Batch Free Can Be Harmful
by: Kim, Daewoo, et al.
Published: (2024)
by: Kim, Daewoo, et al.
Published: (2024)
Batched DGEMMs for scientific codes running on long vector architectures
by: Banchelli, Fabio, et al.
Published: (2025)
by: Banchelli, Fabio, et al.
Published: (2025)
Batch Denoising for AIGC Service Provisioning in Wireless Edge Networks
by: Xu, Jinghang, et al.
Published: (2025)
by: Xu, Jinghang, et al.
Published: (2025)
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
by: Zhao, Yilong, et al.
Published: (2026)
by: Zhao, Yilong, et al.
Published: (2026)
Empowering Distributed Training with Sparsity-driven Data Synchronization
by: Wang, Zhuang, et al.
Published: (2023)
by: Wang, Zhuang, et al.
Published: (2023)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
by: Chen, Qiaoling, et al.
Published: (2023)
by: Chen, Qiaoling, et al.
Published: (2023)
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
by: Ekelund, Jonah, et al.
Published: (2025)
by: Ekelund, Jonah, et al.
Published: (2025)
Herring: Parallel Batch-Order-Fairness on DAG-based Blockchain Consensus
by: Putnik, Marko, et al.
Published: (2026)
by: Putnik, Marko, et al.
Published: (2026)
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
by: Zhu, Juan, et al.
Published: (2026)
by: Zhu, Juan, et al.
Published: (2026)
Similar Items
-
SkyStore: Cost-Optimized Object Storage Across Regions and Clouds
by: Liu, Shu, et al.
Published: (2025) -
PolarStore: High-Performance Data Compression for Large-Scale Cloud-Native Databases
by: Hu, Qingda, et al.
Published: (2025) -
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
by: Shi, Ge, et al.
Published: (2025) -
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
by: Kang, Xueze, et al.
Published: (2025) -
A Reinforcement Learning Based Backfilling Strategy for HPC Batch Jobs
by: Kolker-Hicks, Elliot, et al.
Published: (2024)