Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Ruoyu, He, Weiran, Huang, Weixiao, Zhang, Yangkun, Zhao, Yikai, Pang, Bo, Xu, Xinran, Shan, Yingdi, Wu, Yongwei, Zhang, Mingxing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
by: Qin, Ruoyu, et al.
Published: (2026)
by: Qin, Ruoyu, et al.
Published: (2026)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
by: Qin, Ruoyu, et al.
Published: (2024)
by: Qin, Ruoyu, et al.
Published: (2024)
Efficient Graph-Based Approximate Nearest Neighbor Search Achieving: Low Latency Without Throughput Loss
by: Luo, Jingjia, et al.
Published: (2025)
by: Luo, Jingjia, et al.
Published: (2025)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
by: Ren, Feng, et al.
Published: (2026)
by: Ren, Feng, et al.
Published: (2026)
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
by: Chen, Shaoyuan, et al.
Published: (2024)
by: Chen, Shaoyuan, et al.
Published: (2024)
Seer: Proactive Revenue-Aware Scheduling for Live Streaming Services in Crowdsourced Cloud-Edge Platforms
by: Huang, Shaoyuan, et al.
Published: (2024)
by: Huang, Shaoyuan, et al.
Published: (2024)
Seer: Predictive Runtime Kernel Selection for Irregular Problems
by: Swann, Ryan, et al.
Published: (2024)
by: Swann, Ryan, et al.
Published: (2024)
TrEnv-X: Transparently Share Serverless Execution Environments Across Different Functions and Nodes
by: Huang, Jialiang, et al.
Published: (2025)
by: Huang, Jialiang, et al.
Published: (2025)
A Fast Confirmation Rule (aka Fast Synchronous Finality) for the Ethereum Consensus Protocol
by: Asgaonkar, Aditya, et al.
Published: (2024)
by: Asgaonkar, Aditya, et al.
Published: (2024)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
by: Sun, Xun, et al.
Published: (2026)
by: Sun, Xun, et al.
Published: (2026)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
Hamster: A Fast Synchronous Byzantine Fault Tolerance Protocol
by: Fu, Ximing, et al.
Published: (2024)
by: Fu, Ximing, et al.
Published: (2024)
BFTBrain: Adaptive BFT Consensus with Reinforcement Learning
by: Wu, Chenyuan, et al.
Published: (2024)
by: Wu, Chenyuan, et al.
Published: (2024)
DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training
by: Wang, Yuanqing, et al.
Published: (2026)
by: Wang, Yuanqing, et al.
Published: (2026)
LLM-Enhanced Deep Reinforcement Learning for Task Offloading in Collaborative Edge Computing
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
Efficient Hierarchical Storage Management Framework Empowered by Reinforcement Learning
by: Zhang, Tianru, et al.
Published: (2022)
by: Zhang, Tianru, et al.
Published: (2022)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
by: Chen, Jiu, et al.
Published: (2026)
by: Chen, Jiu, et al.
Published: (2026)
HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
by: He, Yongjun, et al.
Published: (2025)
by: He, Yongjun, et al.
Published: (2025)
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
by: Tang, Zhenheng, et al.
Published: (2025)
by: Tang, Zhenheng, et al.
Published: (2025)
An Advanced Reinforcement Learning Framework for Online Scheduling of Deferrable Workloads in Cloud Computing
by: Dong, Hang, et al.
Published: (2024)
by: Dong, Hang, et al.
Published: (2024)
BootSeer: Analyzing and Mitigating Initialization Bottlenecks in Large-Scale LLM Training
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
by: Wu, Yongtong, et al.
Published: (2026)
by: Wu, Yongtong, et al.
Published: (2026)
A Reinforcement Learning Based Backfilling Strategy for HPC Batch Jobs
by: Kolker-Hicks, Elliot, et al.
Published: (2024)
by: Kolker-Hicks, Elliot, et al.
Published: (2024)
Resource-efficient Parallel Split Learning in Heterogeneous Edge Computing
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
A Reinforcement Learning-Driven Task Scheduling Algorithm for Multi-Tenant Distributed Systems
by: Zhang, Xiaopei, et al.
Published: (2025)
by: Zhang, Xiaopei, et al.
Published: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
by: Wang, Chong, et al.
Published: (2026)
by: Wang, Chong, et al.
Published: (2026)
DFPL: Decentralized Federated Prototype Learning Across Heterogeneous Data Distributions
by: Zhang, Hongliang, et al.
Published: (2025)
by: Zhang, Hongliang, et al.
Published: (2025)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
by: Fan, Jiakun, et al.
Published: (2025)
by: Fan, Jiakun, et al.
Published: (2025)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
by: Li, Haley, et al.
Published: (2026)
by: Li, Haley, et al.
Published: (2026)
HLoRA: Efficient Federated Learning System for LLM Heterogeneous Fine-Tuning
by: Liu, Qianli, et al.
Published: (2025)
by: Liu, Qianli, et al.
Published: (2025)
FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUs
by: Zhang, Haoyuan, et al.
Published: (2025)
by: Zhang, Haoyuan, et al.
Published: (2025)
SAIR: Cost-Efficient Multi-Stage ML Pipeline Autoscaling via In-Context Reinforcement Learning
by: Su, Jianchang, et al.
Published: (2026)
by: Su, Jianchang, et al.
Published: (2026)
Fast Online Digital Twinning on FPGA for Mission Critical Applications
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
SplitLLM: Hierarchical Split Learning for Large Language Model over Wireless Network
by: Zhang, Songge, et al.
Published: (2025)
by: Zhang, Songge, et al.
Published: (2025)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
by: Wu, Siyu, et al.
Published: (2025)
by: Wu, Siyu, et al.
Published: (2025)
FWeb3: A Practical Incentive-Aware Federated Learning Framework
by: Yan, Peishen, et al.
Published: (2026)
by: Yan, Peishen, et al.
Published: (2026)
Fast State Restoration in LLM Serving with HCache
by: Gao, Shiwei, et al.
Published: (2024)
by: Gao, Shiwei, et al.
Published: (2024)
HiRL: Hierarchical Reinforcement Learning for Coordinated Resource Management in Heterogeneous Edge Computing
by: Zhu, Jianyong, et al.
Published: (2026)
by: Zhu, Jianyong, et al.
Published: (2026)
Lodestar: An Online-Learning LLM Inference Router
by: Lim, Gangmuk, et al.
Published: (2026)
by: Lim, Gangmuk, et al.
Published: (2026)
Similar Items
-
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
by: Qin, Ruoyu, et al.
Published: (2026) -
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
by: Qin, Ruoyu, et al.
Published: (2024) -
Efficient Graph-Based Approximate Nearest Neighbor Search Achieving: Low Latency Without Throughput Loss
by: Luo, Jingjia, et al.
Published: (2025) -
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
by: Ren, Feng, et al.
Published: (2026) -
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
by: Chen, Shaoyuan, et al.
Published: (2024)