StageCraft: Execution Aware Mitigation of Distractor and Obstruction Failures in VLA Models
Fuente:
arXiv
Saved in:
| Main Authors: | Pangaonkar, Kartikay Milind, Rath, Prabin, Patil, Omkar, Gopalan, Nakul |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Factorizing Diffusion Policies for Observation Modality Prioritization
by: Patil, Omkar, et al.
Published: (2025)
by: Patil, Omkar, et al.
Published: (2025)
XMoP: Whole-Body Control Policy for Zero-shot Cross-Embodiment Neural Motion Planning
by: Rath, Prabin Kumar, et al.
Published: (2024)
by: Rath, Prabin Kumar, et al.
Published: (2024)
Composing Diffusion Policies for Few-shot Learning of Movement Trajectories
by: Patil, Omkar, et al.
Published: (2024)
by: Patil, Omkar, et al.
Published: (2024)
PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations
by: Gupta, Anmol, et al.
Published: (2026)
by: Gupta, Anmol, et al.
Published: (2026)
Learning Sequential Kinematic Models from Demonstrations for Multi-Jointed Articulated Objects
by: Gupta, Anmol, et al.
Published: (2025)
by: Gupta, Anmol, et al.
Published: (2025)
Meanings and Measurements: Multi-Agent Probabilistic Grounding for Vision-Language Navigation
by: Padhan, Swagat, et al.
Published: (2026)
by: Padhan, Swagat, et al.
Published: (2026)
You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector
by: Patil, Omkar, et al.
Published: (2026)
by: Patil, Omkar, et al.
Published: (2026)
SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models
by: Li, Meng, et al.
Published: (2025)
by: Li, Meng, et al.
Published: (2025)
STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models
by: Xu, Feng, et al.
Published: (2025)
by: Xu, Feng, et al.
Published: (2025)
SeqVLA: Sequential Task Execution for Long-Horizon Manipulation with Completion-Aware Vision-Language-Action Model
by: Yang, Ran, et al.
Published: (2025)
by: Yang, Ran, et al.
Published: (2025)
Green-VLA: Staged Vision-Language-Action Model for Generalist Robots
by: Apanasevich, I., et al.
Published: (2026)
by: Apanasevich, I., et al.
Published: (2026)
Continual Robot Skill and Task Learning via Dialogue
by: Gu, Weiwei, et al.
Published: (2024)
by: Gu, Weiwei, et al.
Published: (2024)
ReconVLA: An Uncertainty-Guided and Failure-Aware Vision-Language-Action Framework for Robotic Control
by: Chen, Lingling, et al.
Published: (2026)
by: Chen, Lingling, et al.
Published: (2026)
GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
by: Abouzeid, Ali, et al.
Published: (2025)
by: Abouzeid, Ali, et al.
Published: (2025)
Long-Term Memory for VLA-based Agents in Open-World Task Execution
by: Huang, Xu, et al.
Published: (2026)
by: Huang, Xu, et al.
Published: (2026)
FUTURE-VLA: Forecasting Unified Trajectories Under Real-time Execution
by: Fan, Jingjing, et al.
Published: (2026)
by: Fan, Jingjing, et al.
Published: (2026)
PAPO-VLA: Planning-Aware Policy Optimization for Vision-Language-Action Models
by: Guo, Peizheng, et al.
Published: (2026)
by: Guo, Peizheng, et al.
Published: (2026)
Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring
by: Park, Seongheon, et al.
Published: (2026)
by: Park, Seongheon, et al.
Published: (2026)
StyleVLA: Driving Style-Aware Vision Language Action Model for Autonomous Driving
by: Gao, Yuan, et al.
Published: (2026)
by: Gao, Yuan, et al.
Published: (2026)
FPC-VLA: A Vision-Language-Action Framework with a Supervisor for Failure Prediction and Correction
by: Yang, Yifan, et al.
Published: (2025)
by: Yang, Yifan, et al.
Published: (2025)
LLM-Craft: Robotic Crafting of Elasto-Plastic Objects with Large Language Models
by: Bartsch, Alison, et al.
Published: (2024)
by: Bartsch, Alison, et al.
Published: (2024)
Selective Perception for Robot: Task-Aware Attention in Multimodal VLA
by: Son, Young-Chae, et al.
Published: (2026)
by: Son, Young-Chae, et al.
Published: (2026)
DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution
by: Yue, Yang, et al.
Published: (2024)
by: Yue, Yang, et al.
Published: (2024)
VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
A Factory-Floor Deployment Case Study of VLA Pipelines for Industrial Packaging Task: Workflow, Failures, and Lessons
by: Zhu, Brian, et al.
Published: (2026)
by: Zhu, Brian, et al.
Published: (2026)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
by: Gu, Chenyang, et al.
Published: (2025)
by: Gu, Chenyang, et al.
Published: (2025)
RadarSFD: Single-Frame Diffusion with Pretrained Priors for Radar Point Clouds
by: Zhao, Bin, et al.
Published: (2025)
by: Zhao, Bin, et al.
Published: (2025)
ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2025)
by: Van Vo, Tuan, et al.
Published: (2025)
How Fast Can I Run My VLA? Demystifying VLA Inference Performance with VLA-Perf
by: Jiang, Wenqi, et al.
Published: (2026)
by: Jiang, Wenqi, et al.
Published: (2026)
DEFLECT: Delay-Robust Execution via Flow-matching Likelihood-Estimated Counterfactual Tuning for VLA Policies
by: Zhu, Yixiang, et al.
Published: (2026)
by: Zhu, Yixiang, et al.
Published: (2026)
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
ProgVLA: Progress-Aware Robot Manipulation Skill Learning
by: Kim, Seungsu, et al.
Published: (2026)
by: Kim, Seungsu, et al.
Published: (2026)
ST-VLA: Enabling 4D-Aware Spatiotemporal Understanding for General Robot Manipulation
by: Wu, You, et al.
Published: (2026)
by: Wu, You, et al.
Published: (2026)
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
by: Zhang, Kaidi, et al.
Published: (2026)
by: Zhang, Kaidi, et al.
Published: (2026)
Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
by: Feng, Yicheng, et al.
Published: (2025)
by: Feng, Yicheng, et al.
Published: (2025)
BlockVLA: Accelerating Autoregressive VLA via Block Diffusion Finetuning
by: Wang, Ruiheng, et al.
Published: (2026)
by: Wang, Ruiheng, et al.
Published: (2026)
EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
by: Lin, Min, et al.
Published: (2025)
by: Lin, Min, et al.
Published: (2025)
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Similar Items
-
Factorizing Diffusion Policies for Observation Modality Prioritization
by: Patil, Omkar, et al.
Published: (2025) -
XMoP: Whole-Body Control Policy for Zero-shot Cross-Embodiment Neural Motion Planning
by: Rath, Prabin Kumar, et al.
Published: (2024) -
Composing Diffusion Policies for Few-shot Learning of Movement Trajectories
by: Patil, Omkar, et al.
Published: (2024) -
PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations
by: Gupta, Anmol, et al.
Published: (2026) -
Learning Sequential Kinematic Models from Demonstrations for Multi-Jointed Articulated Objects
by: Gupta, Anmol, et al.
Published: (2025)