SOP: A Scalable Online Post-Training System for Vision-Language-Action Models
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Mingjie, Feng, Siyuan, Zhang, Qinglin, Li, Xinchen, Song, Jianheng, Qu, Chendi, Wang, Yi, Li, Chuankang, Xiong, Ziyu, Chen, Zhi, Liu, Yi, Luo, Jianlan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
FM SO.P: A Progressive Task Mixture Framework with Automatic Evaluation for Cross-Domain SOP Understanding
by: Huang, Siyuan, et al.
Published: (2026)
by: Huang, Siyuan, et al.
Published: (2026)
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
by: Xu, Siyuan, et al.
Published: (2026)
by: Xu, Siyuan, et al.
Published: (2026)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
by: Zhong, Linqing, et al.
Published: (2026)
by: Zhong, Linqing, et al.
Published: (2026)
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
by: Liu, Zhi
Published: (2026)
by: Liu, Zhi
Published: (2026)
ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents
by: Li, Zhigen, et al.
Published: (2024)
by: Li, Zhigen, et al.
Published: (2024)
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
by: Yang, Rushuai, et al.
Published: (2026)
by: Yang, Rushuai, et al.
Published: (2026)
Genie Centurion: Accelerating Scalable Real-World Robot Training with Human Rewind-and-Refine Guidance
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
QMamba: Post-Training Quantization for Vision State Space Models
by: Li, Yinglong, et al.
Published: (2025)
by: Li, Yinglong, et al.
Published: (2025)
Calibration of Parameters of the Equation of State of Black Powder Based on Underwater Explosion Experiments and Genetic Algorithm
by: Yuzhu Zhang, et al.
Published: (2025)
by: Yuzhu Zhang, et al.
Published: (2025)
Post-Training Quantization for Vision Mamba with k-Scaled Quantization and Reparameterization
by: Shi, Bo-Yun, et al.
Published: (2025)
by: Shi, Bo-Yun, et al.
Published: (2025)
Interactive Post-Training for Vision-Language-Action Models
by: Tan, Shuhan, et al.
Published: (2025)
by: Tan, Shuhan, et al.
Published: (2025)
Flow-of-Action: SOP Enhanced LLM-Based Multi-Agent System for Root Cause Analysis
by: Pei, Changhua, et al.
Published: (2025)
by: Pei, Changhua, et al.
Published: (2025)
Strategic Cheating in Young Children
by: Li Zhao, et al.
Published: (2025)
by: Li Zhao, et al.
Published: (2025)
Optimal Unpredictable Control for Linear Systems
by: Qu, Chendi, et al.
Published: (2025)
by: Qu, Chendi, et al.
Published: (2025)
Event-Grounded Sparse Autoencoders for Vision-Language-Action Policies
by: Jin, Xinchen, et al.
Published: (2026)
by: Jin, Xinchen, et al.
Published: (2026)
RLIF: Interactive Imitation Learning as Reinforcement Learning
by: Luo, Jianlan, et al.
Published: (2023)
by: Luo, Jianlan, et al.
Published: (2023)
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation
by: Feng, Yunhai, et al.
Published: (2025)
by: Feng, Yunhai, et al.
Published: (2025)
WOBEC SOP Echosounder
by: Flores, Hauke, et al.
Published: (2025)
by: Flores, Hauke, et al.
Published: (2025)
3DIOC: Direct Data-Driven Inverse Optimal Control for LTI Systems
by: Qu, Chendi, et al.
Published: (2024)
by: Qu, Chendi, et al.
Published: (2024)
HQViT: Hybrid Quantum Vision Transformer for Image Classification
by: Zhang, Hui, et al.
Published: (2025)
by: Zhang, Hui, et al.
Published: (2025)
Vision SmolMamba: Spike-Guided Token Pruning for Energy-Efficient Spiking State-Space Vision Models
by: Bai, Dewei, et al.
Published: (2026)
by: Bai, Dewei, et al.
Published: (2026)
Introducing Multimodal Paradigm for Learning Sleep Staging PSG via General-Purpose Model
by: Zhou, Jianheng, et al.
Published: (2025)
by: Zhou, Jianheng, et al.
Published: (2025)
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
by: Li, Mingjie, et al.
Published: (2026)
by: Li, Mingjie, et al.
Published: (2026)
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
by: Xu, Charles, et al.
Published: (2024)
by: Xu, Charles, et al.
Published: (2024)
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
by: Yang, Siyuan, et al.
Published: (2025)
by: Yang, Siyuan, et al.
Published: (2025)
IIP-Mixer:Intra-Inter Patch Mixing Architecture for Battery Remaining Useful Life Prediction
by: Ye, Guangzai, et al.
Published: (2024)
by: Ye, Guangzai, et al.
Published: (2024)
DORA: Dynamic Online Reinforcement Agent for Token Merging in Vision Transformers
by: He, Kaixuan, et al.
Published: (2026)
by: He, Kaixuan, et al.
Published: (2026)
NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models
by: Zhu, Ziyue, et al.
Published: (2026)
by: Zhu, Ziyue, et al.
Published: (2026)
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models
by: Fang, Zhou, et al.
Published: (2026)
by: Fang, Zhou, et al.
Published: (2026)
Topic mining based on fine-tuning Sentence-BERT and LDA
by: Li, Jianheng, et al.
Published: (2025)
by: Li, Jianheng, et al.
Published: (2025)
Successful Treatment of Intra‐Abdominal Carbapenem‐Resistant Acinetobacter baumannii Infection With Sulbactam–Durlobactam in a Child With Acute Liver Failure Following Auxiliary Liver Transplantation
by: Hao Feng Xiong, et al.
Published: (2026)
by: Hao Feng Xiong, et al.
Published: (2026)
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
by: Yu, Hanxun, et al.
Published: (2026)
by: Yu, Hanxun, et al.
Published: (2026)
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
by: Wang, Zhixin, et al.
Published: (2025)
by: Wang, Zhixin, et al.
Published: (2025)
Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model
by: Chen, Yi-Chia, et al.
Published: (2024)
by: Chen, Yi-Chia, et al.
Published: (2024)
Actor-Critic based Online Data Mixing For Language Model Pre-Training
by: Ma, Jing, et al.
Published: (2025)
by: Ma, Jing, et al.
Published: (2025)
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
by: Li, Wanli, et al.
Published: (2026)
by: Li, Wanli, et al.
Published: (2026)
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Similar Items
-
Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
by: Wang, Yi, et al.
Published: (2026) -
FM SO.P: A Progressive Task Mixture Framework with Automatic Evaluation for Cross-Domain SOP Understanding
by: Huang, Siyuan, et al.
Published: (2026) -
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
by: Liu, Yi, et al.
Published: (2025) -
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
by: Xu, Siyuan, et al.
Published: (2026) -
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
by: Zhong, Linqing, et al.
Published: (2026)