SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fei, Senyu, Wang, Siyin, Ji, Li, Li, Ao, Zhang, Shiduo, Liu, Liming, Hou, Jinlong, Gong, Jingjing, Zhao, Xianzhong, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
von: Fei, Senyu, et al.
Veröffentlicht: (2025)
von: Fei, Senyu, et al.
Veröffentlicht: (2025)
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
von: Shi, Junhao, et al.
Veröffentlicht: (2025)
von: Shi, Junhao, et al.
Veröffentlicht: (2025)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
World Action Models: The Next Frontier in Embodied AI
von: Wang, Siyin, et al.
Veröffentlicht: (2026)
von: Wang, Siyin, et al.
Veröffentlicht: (2026)
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
von: Fei, Zhaoye, et al.
Veröffentlicht: (2025)
von: Fei, Zhaoye, et al.
Veröffentlicht: (2025)
ActionCodec: What Makes for Good Action Tokenizers
von: Dong, Zibin, et al.
Veröffentlicht: (2026)
von: Dong, Zibin, et al.
Veröffentlicht: (2026)
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
von: Zhang, Shiduo, et al.
Veröffentlicht: (2024)
von: Zhang, Shiduo, et al.
Veröffentlicht: (2024)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
von: Zhao, Chen, et al.
Veröffentlicht: (2026)
von: Zhao, Chen, et al.
Veröffentlicht: (2026)
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
von: Hou, Zhi, et al.
Veröffentlicht: (2025)
von: Hou, Zhi, et al.
Veröffentlicht: (2025)
Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
von: Li, Pengxiang, et al.
Veröffentlicht: (2025)
von: Li, Pengxiang, et al.
Veröffentlicht: (2025)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
von: Ye, Wencheng, et al.
Veröffentlicht: (2025)
von: Ye, Wencheng, et al.
Veröffentlicht: (2025)
Unified Vision-Language-Action Model
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
von: Lv, Qi, et al.
Veröffentlicht: (2025)
von: Lv, Qi, et al.
Veröffentlicht: (2025)
LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models
von: Saxena, Pranav, et al.
Veröffentlicht: (2025)
von: Saxena, Pranav, et al.
Veröffentlicht: (2025)
Conceptual and Design Principles for a Self-Referential Algorithm Mimicking Neuronal Assembly Functions
von: Totaro, Paolo, et al.
Veröffentlicht: (2025)
von: Totaro, Paolo, et al.
Veröffentlicht: (2025)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
von: Liang, Zhixuan, et al.
Veröffentlicht: (2025)
von: Liang, Zhixuan, et al.
Veröffentlicht: (2025)
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
von: Fu, Yiyang, et al.
Veröffentlicht: (2026)
von: Fu, Yiyang, et al.
Veröffentlicht: (2026)
Stable Language Guidance for Vision-Language-Action Models
von: Zhan, Zhihao, et al.
Veröffentlicht: (2026)
von: Zhan, Zhihao, et al.
Veröffentlicht: (2026)
LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action Models
von: Hou, Yuchen, et al.
Veröffentlicht: (2026)
von: Hou, Yuchen, et al.
Veröffentlicht: (2026)
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
von: Lin, Tao, et al.
Veröffentlicht: (2025)
von: Lin, Tao, et al.
Veröffentlicht: (2025)
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
von: Zhang, Borong, et al.
Veröffentlicht: (2025)
von: Zhang, Borong, et al.
Veröffentlicht: (2025)
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
von: Lin, Juyi, et al.
Veröffentlicht: (2025)
von: Lin, Juyi, et al.
Veröffentlicht: (2025)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
von: Xiao, Lei, et al.
Veröffentlicht: (2025)
von: Xiao, Lei, et al.
Veröffentlicht: (2025)
Joint Action Language Modelling for Transparent Policy Execution
von: Wulff, Theodor, et al.
Veröffentlicht: (2025)
von: Wulff, Theodor, et al.
Veröffentlicht: (2025)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026)
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026)
HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
von: Koo, Myungkyu, et al.
Veröffentlicht: (2025)
von: Koo, Myungkyu, et al.
Veröffentlicht: (2025)
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
von: Xu, Kechun, et al.
Veröffentlicht: (2025)
von: Xu, Kechun, et al.
Veröffentlicht: (2025)
HarvestFlex: Strawberry Harvesting via Vision-Language-Action Policy Adaptation in the Wild
von: Zhao, Ziyang, et al.
Veröffentlicht: (2026)
von: Zhao, Ziyang, et al.
Veröffentlicht: (2026)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
von: Lin, Yihan, et al.
Veröffentlicht: (2026)
von: Lin, Yihan, et al.
Veröffentlicht: (2026)
Efficient Long-Horizon Vision-Language-Action Models via Static-Dynamic Disentanglement
von: Qiu, Weikang, et al.
Veröffentlicht: (2026)
von: Qiu, Weikang, et al.
Veröffentlicht: (2026)
EdgeVLA: Efficient Vision-Language-Action Models
von: Budzianowski, Paweł, et al.
Veröffentlicht: (2025)
von: Budzianowski, Paweł, et al.
Veröffentlicht: (2025)
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
von: Wang, Hanzhen, et al.
Veröffentlicht: (2025)
von: Wang, Hanzhen, et al.
Veröffentlicht: (2025)
QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization
von: Xu, Yuhao, et al.
Veröffentlicht: (2026)
von: Xu, Yuhao, et al.
Veröffentlicht: (2026)
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
von: Fei, Senyu, et al.
Veröffentlicht: (2025) -
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
von: Wang, Siyin, et al.
Veröffentlicht: (2025) -
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
von: Shi, Junhao, et al.
Veröffentlicht: (2025) -
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
von: Liu, Yicheng, et al.
Veröffentlicht: (2025) -
World Action Models: The Next Frontier in Embodied AI
von: Wang, Siyin, et al.
Veröffentlicht: (2026)