Improving Robustness of Vision-Language-Action Models by Restoring Corrupted Visual Inputs
Fuente:
arXiv
Salvato in:
| Autori principali: | Orjuela, Daniel Yezid Guarnizo, Scappatura, Leonardo, Di Gennaro, Veronica, Izzo, Riccardo Andrea, Bardaro, Gianluca, Matteucci, Matteo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models
di: Izzo, Riccardo Andrea, et al.
Pubblicazione: (2026)
di: Izzo, Riccardo Andrea, et al.
Pubblicazione: (2026)
BTGenBot-2: Efficient Behavior Tree Generation with Small Language Models
di: Izzo, Riccardo Andrea, et al.
Pubblicazione: (2026)
di: Izzo, Riccardo Andrea, et al.
Pubblicazione: (2026)
Multimodal Behavior Tree Generation: A Small Vision-Language Model for Robot Task Planning
di: Battistini, Cristiano, et al.
Pubblicazione: (2026)
di: Battistini, Cristiano, et al.
Pubblicazione: (2026)
BTGenBot: Behavior Tree Generation for Robotic Tasks with Lightweight LLMs
di: Izzo, Riccardo Andrea, et al.
Pubblicazione: (2024)
di: Izzo, Riccardo Andrea, et al.
Pubblicazione: (2024)
A Spatio-temporal Graph Network Allowing Incomplete Trajectory Input for Pedestrian Trajectory Prediction
di: Long, Juncen, et al.
Pubblicazione: (2025)
di: Long, Juncen, et al.
Pubblicazione: (2025)
The Empirical Impact of Forgetting and Transfer in Continual Visual Odometry
di: Cudrano, Paolo, et al.
Pubblicazione: (2024)
di: Cudrano, Paolo, et al.
Pubblicazione: (2024)
Advancements in Radar Odometry
di: Frosi, Matteo, et al.
Pubblicazione: (2023)
di: Frosi, Matteo, et al.
Pubblicazione: (2023)
Enhancing Agricultural Environment Perception via Active Vision and Zero-Shot Learning
di: La Greca, Michele Carlo, et al.
Pubblicazione: (2024)
di: La Greca, Michele Carlo, et al.
Pubblicazione: (2024)
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
di: Li, Jinming, et al.
Pubblicazione: (2024)
di: Li, Jinming, et al.
Pubblicazione: (2024)
IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks
di: Hannus, Eric, et al.
Pubblicazione: (2025)
di: Hannus, Eric, et al.
Pubblicazione: (2025)
Run-time Observation Interventions Make Vision-Language-Action Models More Visually Robust
di: Hancock, Asher J., et al.
Pubblicazione: (2024)
di: Hancock, Asher J., et al.
Pubblicazione: (2024)
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
di: Zhang, Yichi, et al.
Pubblicazione: (2026)
di: Zhang, Yichi, et al.
Pubblicazione: (2026)
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
di: Wang, Zixuan, et al.
Pubblicazione: (2026)
di: Wang, Zixuan, et al.
Pubblicazione: (2026)
Enhancing Robot Explanation Capabilities through Vision-Language Models: a Preliminary Study by Interpreting Visual Inputs for Improved Human-Robot Interaction
di: Sobrín-Hidalgo, David, et al.
Pubblicazione: (2024)
di: Sobrín-Hidalgo, David, et al.
Pubblicazione: (2024)
Evaluating and Improving the Robustness of LiDAR Odometry and Localization Under Real-World Corruptions
di: Yang, Bo, et al.
Pubblicazione: (2024)
di: Yang, Bo, et al.
Pubblicazione: (2024)
EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents
di: Riva, Paolo, et al.
Pubblicazione: (2026)
di: Riva, Paolo, et al.
Pubblicazione: (2026)
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
di: Xiao, Lei, et al.
Pubblicazione: (2025)
di: Xiao, Lei, et al.
Pubblicazione: (2025)
VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
di: Chen, Guanyan, et al.
Pubblicazione: (2024)
di: Chen, Guanyan, et al.
Pubblicazione: (2024)
RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models
di: Luo, Jingzhou, et al.
Pubblicazione: (2026)
di: Luo, Jingzhou, et al.
Pubblicazione: (2026)
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
di: Glossop, Catherine, et al.
Pubblicazione: (2025)
di: Glossop, Catherine, et al.
Pubblicazione: (2025)
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving
di: Zhang, Liangdong, et al.
Pubblicazione: (2026)
di: Zhang, Liangdong, et al.
Pubblicazione: (2026)
Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching
di: Wei, Yujie, et al.
Pubblicazione: (2026)
di: Wei, Yujie, et al.
Pubblicazione: (2026)
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
di: Zhang, Kaidi, et al.
Pubblicazione: (2026)
di: Zhang, Kaidi, et al.
Pubblicazione: (2026)
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
di: Vo, Khoa, et al.
Pubblicazione: (2025)
di: Vo, Khoa, et al.
Pubblicazione: (2025)
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
di: Zhong, Zhide, et al.
Pubblicazione: (2025)
di: Zhong, Zhide, et al.
Pubblicazione: (2025)
STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations
di: Xie, Yuhan, et al.
Pubblicazione: (2026)
di: Xie, Yuhan, et al.
Pubblicazione: (2026)
Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging
di: Yadav, Yajat, et al.
Pubblicazione: (2025)
di: Yadav, Yajat, et al.
Pubblicazione: (2025)
Improving Pre-Trained Vision-Language-Action Policies with Model-Based Search
di: Neary, Cyrus, et al.
Pubblicazione: (2025)
di: Neary, Cyrus, et al.
Pubblicazione: (2025)
Concept-Based Dictionary Learning for Inference-Time Safety in Vision Language Action Models
di: Wen, Siqi, et al.
Pubblicazione: (2026)
di: Wen, Siqi, et al.
Pubblicazione: (2026)
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
di: Li, Pengteng, et al.
Pubblicazione: (2026)
di: Li, Pengteng, et al.
Pubblicazione: (2026)
Grounding Hierarchical Vision-Language-Action Models Through Explicit Language-Action Alignment
di: Wulff, Theodor, et al.
Pubblicazione: (2026)
di: Wulff, Theodor, et al.
Pubblicazione: (2026)
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
di: Liang, Yuanchang, et al.
Pubblicazione: (2026)
di: Liang, Yuanchang, et al.
Pubblicazione: (2026)
Boosting Vision-Language-Action Finetuning with Feasible Action Neighborhood Prior
di: Niu, Haochen, et al.
Pubblicazione: (2026)
di: Niu, Haochen, et al.
Pubblicazione: (2026)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
di: Zhong, Yifan, et al.
Pubblicazione: (2025)
di: Zhong, Yifan, et al.
Pubblicazione: (2025)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
di: Zhong, Linqing, et al.
Pubblicazione: (2026)
di: Zhong, Linqing, et al.
Pubblicazione: (2026)
AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models
di: Heo, Hyeongjun, et al.
Pubblicazione: (2026)
di: Heo, Hyeongjun, et al.
Pubblicazione: (2026)
Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
di: Lee, Sangoh, et al.
Pubblicazione: (2025)
di: Lee, Sangoh, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models
di: Izzo, Riccardo Andrea, et al.
Pubblicazione: (2026) -
BTGenBot-2: Efficient Behavior Tree Generation with Small Language Models
di: Izzo, Riccardo Andrea, et al.
Pubblicazione: (2026) -
Multimodal Behavior Tree Generation: A Small Vision-Language Model for Robot Task Planning
di: Battistini, Cristiano, et al.
Pubblicazione: (2026) -
BTGenBot: Behavior Tree Generation for Robotic Tasks with Lightweight LLMs
di: Izzo, Riccardo Andrea, et al.
Pubblicazione: (2024) -
A Spatio-temporal Graph Network Allowing Incomplete Trajectory Input for Pedestrian Trajectory Prediction
di: Long, Juncen, et al.
Pubblicazione: (2025)