Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Kwok, Jacky, Zhang, Xilun, Xu, Mengdi, Liu, Yuejiang, Mirhoseini, Azalia, Finn, Chelsea, Pavone, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
by: Kwok, Jacky, et al.
Published: (2025)
by: Kwok, Jacky, et al.
Published: (2025)
Self-Guided Action Diffusion
by: Malhotra, Rhea, et al.
Published: (2025)
by: Malhotra, Rhea, et al.
Published: (2025)
Learning Long-Context Diffusion Policies via Past-Token Prediction
by: Torne, Marcel, et al.
Published: (2025)
by: Torne, Marcel, et al.
Published: (2025)
VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model
by: Guo, Yanjiang, et al.
Published: (2026)
by: Guo, Yanjiang, et al.
Published: (2026)
Convex Hulls of Reachable Sets
by: Lew, Thomas, et al.
Published: (2023)
by: Lew, Thomas, et al.
Published: (2023)
Observing and Controlling Features in Vision-Language-Action Models
by: Buurmeijer, Hugo, et al.
Published: (2026)
by: Buurmeijer, Hugo, et al.
Published: (2026)
Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling
by: Liu, Yuejiang, et al.
Published: (2024)
by: Liu, Yuejiang, et al.
Published: (2024)
Curating Demonstrations using Online Experience
by: Chen, Annie S., et al.
Published: (2025)
by: Chen, Annie S., et al.
Published: (2025)
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
by: Torne, Marcel, et al.
Published: (2026)
by: Torne, Marcel, et al.
Published: (2026)
VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer
by: Hu, Songqiao, et al.
Published: (2025)
by: Hu, Songqiao, et al.
Published: (2025)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
by: Kim, Moo Jin, et al.
Published: (2025)
by: Kim, Moo Jin, et al.
Published: (2025)
Constraint-Aware Reinforcement Learning via Adaptive Action Scaling
by: Dawood, Murad, et al.
Published: (2025)
by: Dawood, Murad, et al.
Published: (2025)
Real-Time Anomaly Detection and Reactive Planning with Large Language Models
by: Sinha, Rohan, et al.
Published: (2024)
by: Sinha, Rohan, et al.
Published: (2024)
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
by: Wang, Yihao, et al.
Published: (2025)
by: Wang, Yihao, et al.
Published: (2025)
Towards Unified Probabilistic Verification and Validation of Vision-Based Autonomy
by: Peper, Jordan, et al.
Published: (2025)
by: Peper, Jordan, et al.
Published: (2025)
Emergence of Human to Robot Transfer in Vision-Language-Action Models
by: Kareer, Simar, et al.
Published: (2025)
by: Kareer, Simar, et al.
Published: (2025)
Joint Optimization of Autonomous Electric Vehicle Fleet Operations and Charging Station Siting
by: Luke, Justin, et al.
Published: (2021)
by: Luke, Justin, et al.
Published: (2021)
EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
Multi-Timescale Model Predictive Control for Slow-Fast Systems
by: Schroth, Lukas, et al.
Published: (2025)
by: Schroth, Lukas, et al.
Published: (2025)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
Agile Tradespace Exploration for Space Rendezvous Mission Design via Transformers
by: Takubo, Yuji, et al.
Published: (2025)
by: Takubo, Yuji, et al.
Published: (2025)
Large Scale Robotic Material Handling: Learning, Planning, and Control
by: Spinelli, Filippo A., et al.
Published: (2025)
by: Spinelli, Filippo A., et al.
Published: (2025)
Language-Conditioned Safe Trajectory Generation for Spacecraft Rendezvous
by: Takubo, Yuji, et al.
Published: (2025)
by: Takubo, Yuji, et al.
Published: (2025)
Vision-Language-Action Models for Selective Robotic Disassembly: A Case Study on Critical Component Extraction from Desktops
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
Taming High-Dimensional Dynamics: Learning Optimal Projections onto Spectral Submanifolds
by: Buurmeijer, Hugo, et al.
Published: (2025)
by: Buurmeijer, Hugo, et al.
Published: (2025)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
by: Hou, Zhi, et al.
Published: (2025)
by: Hou, Zhi, et al.
Published: (2025)
Perfecting Periodic Trajectory Tracking: Model Predictive Control with a Periodic Observer ($Π$-MPC)
by: Pabon, Luis, et al.
Published: (2024)
by: Pabon, Luis, et al.
Published: (2024)
Realistic Extreme Behavior Generation for Improved AV Testing
by: Dyro, Robert, et al.
Published: (2024)
by: Dyro, Robert, et al.
Published: (2024)
Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
by: Fu, Zipeng, et al.
Published: (2024)
by: Fu, Zipeng, et al.
Published: (2024)
Towards Robust Spacecraft Trajectory Optimization via Transformers
by: Takubo, Yuji, et al.
Published: (2024)
by: Takubo, Yuji, et al.
Published: (2024)
Nipping the Drift in the Bud: Retrospective Rectification for Robust Vision-Language Navigation
by: He, Gang, et al.
Published: (2026)
by: He, Gang, et al.
Published: (2026)
World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry
by: Liu, Yuejiang, et al.
Published: (2026)
by: Liu, Yuejiang, et al.
Published: (2026)
Reachability-Aware Time Scaling for Path Tracking
by: Gholampour, Hossein, et al.
Published: (2026)
by: Gholampour, Hossein, et al.
Published: (2026)
On the Role of Temperature Sampling in Test-Time Scaling
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
PerFACT: Motion Policy with LLM-Powered Dataset Synthesis and Fusion Action-Chunking Transformers
by: Soleymanzadeh, Davood, et al.
Published: (2025)
by: Soleymanzadeh, Davood, et al.
Published: (2025)
Real-time Control of Electric Autonomous Mobility-on-Demand Systems via Graph Reinforcement Learning
by: Singhal, Aaryan, et al.
Published: (2023)
by: Singhal, Aaryan, et al.
Published: (2023)
Safety Evaluation of Motion Plans Using Trajectory Predictors as Forward Reachable Set Estimators
by: Chakraborty, Kaustav, et al.
Published: (2025)
by: Chakraborty, Kaustav, et al.
Published: (2025)
PaCo-VLA: Passivity-Shielded Compliance Prior for Contact-Rich Vision-Language-Action Manipulation
by: Cao, Haofan, et al.
Published: (2026)
by: Cao, Haofan, et al.
Published: (2026)
SignVLA: A Gloss-Free Vision-Language-Action Framework for Real-Time Sign Language-Guided Robotic Manipulation
by: Tan, Xinyu, et al.
Published: (2026)
by: Tan, Xinyu, et al.
Published: (2026)
Similar Items
-
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
by: Kwok, Jacky, et al.
Published: (2025) -
Self-Guided Action Diffusion
by: Malhotra, Rhea, et al.
Published: (2025) -
Learning Long-Context Diffusion Policies via Past-Token Prediction
by: Torne, Marcel, et al.
Published: (2025) -
VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model
by: Guo, Yanjiang, et al.
Published: (2026) -
Convex Hulls of Reachable Sets
by: Lew, Thomas, et al.
Published: (2023)