VISTA: Enhancing Visual Conditioning via Track-Following Preference Optimization in Vision-Language-Action Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Yiye, Jian, Yanan, Dong, Xiaoyi, Cao, Shuxin, Wu, Jing, Vela, Patricio, Lundell, Benjamin E., Chen, Dongdong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Schema-Guided Scene-Graph Reasoning based on Multi-Agent Large Language Model System
por: Chen, Yiye, et al.
Publicado: (2025)
por: Chen, Yiye, et al.
Publicado: (2025)
Good Weights: Proactive, Adaptive Dead Reckoning Fusion for Continuous and Robust Visual SLAM
por: Du, Yanwei, et al.
Publicado: (2025)
por: Du, Yanwei, et al.
Publicado: (2025)
VISTA: Generative Visual Imagination for Vision-and-Language Navigation
por: Huang, Yanjia, et al.
Publicado: (2025)
por: Huang, Yanjia, et al.
Publicado: (2025)
VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation
por: Zhang, Chaofan, et al.
Publicado: (2025)
por: Zhang, Chaofan, et al.
Publicado: (2025)
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
por: Glossop, Catherine, et al.
Publicado: (2025)
por: Glossop, Catherine, et al.
Publicado: (2025)
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
por: Wang, Zixuan, et al.
Publicado: (2026)
por: Wang, Zixuan, et al.
Publicado: (2026)
Efficient Iterative Proximal Variational Inference Motion Planning
por: Chang, Zinuo, et al.
Publicado: (2024)
por: Chang, Zinuo, et al.
Publicado: (2024)
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
por: Bai, Shuanghao, et al.
Publicado: (2026)
por: Bai, Shuanghao, et al.
Publicado: (2026)
DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization
por: Lin, Sixu, et al.
Publicado: (2026)
por: Lin, Sixu, et al.
Publicado: (2026)
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
por: Hu, Yucheng, et al.
Publicado: (2026)
por: Hu, Yucheng, et al.
Publicado: (2026)
VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
por: Chen, Guanyan, et al.
Publicado: (2024)
por: Chen, Guanyan, et al.
Publicado: (2024)
GeomPrompt: Geometric Prompt Learning for RGB-D Semantic Segmentation Under Missing and Degraded Depth
por: Jaganathan, Krishna, et al.
Publicado: (2026)
por: Jaganathan, Krishna, et al.
Publicado: (2026)
Point Tracking Improves World Action Models
por: Guan, Jiarui, et al.
Publicado: (2026)
por: Guan, Jiarui, et al.
Publicado: (2026)
A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving
por: Zhang, Liangdong, et al.
Publicado: (2026)
por: Zhang, Liangdong, et al.
Publicado: (2026)
Preference-Conditioned Multi-Objective RL for Integrated Command Tracking and Force Compliance in Humanoid Locomotion
por: Leng, Tingxuan, et al.
Publicado: (2025)
por: Leng, Tingxuan, et al.
Publicado: (2025)
HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning
por: Shou, Quanxin, et al.
Publicado: (2026)
por: Shou, Quanxin, et al.
Publicado: (2026)
QuadPiPS: A Perception-informed Footstep Planner for Quadrupeds With Semantic Affordance Prediction
por: Asselmeier, Max, et al.
Publicado: (2024)
por: Asselmeier, Max, et al.
Publicado: (2024)
Factor Graph-Based Shape Estimation for Continuum Robots via Magnus Expansion
por: Ticozzi, Lorenzo, et al.
Publicado: (2026)
por: Ticozzi, Lorenzo, et al.
Publicado: (2026)
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
por: Chen, Xiaoyu, et al.
Publicado: (2025)
por: Chen, Xiaoyu, et al.
Publicado: (2025)
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
por: Deng, Shengliang, et al.
Publicado: (2025)
por: Deng, Shengliang, et al.
Publicado: (2025)
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
por: Zhong, Zhide, et al.
Publicado: (2025)
por: Zhong, Zhide, et al.
Publicado: (2025)
AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models
por: Li, Xiaoqi, et al.
Publicado: (2026)
por: Li, Xiaoqi, et al.
Publicado: (2026)
MMDVS-LF: Multi-Modal Dynamic Vision Sensor and Eye-Tracking Dataset for Line Following
por: Resch, Felix, et al.
Publicado: (2024)
por: Resch, Felix, et al.
Publicado: (2024)
SURESTEP: An Uncertainty-Aware Trajectory Optimization Framework to Enhance Visual Tool Tracking for Robust Surgical Automation
por: Shinde, Nikhil U., et al.
Publicado: (2024)
por: Shinde, Nikhil U., et al.
Publicado: (2024)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
por: Zhong, Yifan, et al.
Publicado: (2025)
por: Zhong, Yifan, et al.
Publicado: (2025)
Health-Conditioned Vision-Language-Action Models for Malfunction-Aware Robot Control
por: Arslan, Hüseyin, et al.
Publicado: (2026)
por: Arslan, Hüseyin, et al.
Publicado: (2026)
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
por: Zhang, Yichi, et al.
Publicado: (2026)
por: Zhang, Yichi, et al.
Publicado: (2026)
CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization
por: Wang, Xiangyue, et al.
Publicado: (2026)
por: Wang, Xiangyue, et al.
Publicado: (2026)
Human-assisted Robotic Policy Refinement via Action Preference Optimization
por: Xia, Wenke, et al.
Publicado: (2025)
por: Xia, Wenke, et al.
Publicado: (2025)
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
por: Liang, Yuanchang, et al.
Publicado: (2026)
por: Liang, Yuanchang, et al.
Publicado: (2026)
Personalization in Human-Robot Interaction through Preference-based Action Representation Learning
por: Wang, Ruiqi, et al.
Publicado: (2024)
por: Wang, Ruiqi, et al.
Publicado: (2024)
CaFe-TeleVision: A Coarse-to-Fine Teleoperation System with Immersive Situated Visualization for Enhanced Ergonomics
por: Tang, Zixin, et al.
Publicado: (2025)
por: Tang, Zixin, et al.
Publicado: (2025)
DDGC: Generative Deep Dexterous Grasping in Clutter
por: Lundell, Jens, et al.
Publicado: (2021)
por: Lundell, Jens, et al.
Publicado: (2021)
LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments
por: Ding, Hongyu, et al.
Publicado: (2025)
por: Ding, Hongyu, et al.
Publicado: (2025)
Failing Forward: Adaptive Failure-Informed Learning for Vision-Language-Action Models
por: Zheng, Meng, et al.
Publicado: (2026)
por: Zheng, Meng, et al.
Publicado: (2026)
Embodied Learning of Reward for Musculoskeletal Control with Vision Language Models
por: Soedarmadji, Saraswati, et al.
Publicado: (2025)
por: Soedarmadji, Saraswati, et al.
Publicado: (2025)
UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models
por: Zhang, Qiyao, et al.
Publicado: (2026)
por: Zhang, Qiyao, et al.
Publicado: (2026)
Pushing Everything Everywhere All At Once: Probabilistic Prehensile Pushing
por: Perugini, Patrizio, et al.
Publicado: (2025)
por: Perugini, Patrizio, et al.
Publicado: (2025)
CAPGrasp: An $\mathbb{R}^3\times \text{SO(2)-equivariant}$ Continuous Approach-Constrained Generative Grasp Sampler
por: Weng, Zehang, et al.
Publicado: (2023)
por: Weng, Zehang, et al.
Publicado: (2023)
Ejemplares similares
-
Schema-Guided Scene-Graph Reasoning based on Multi-Agent Large Language Model System
por: Chen, Yiye, et al.
Publicado: (2025) -
Good Weights: Proactive, Adaptive Dead Reckoning Fusion for Continuous and Robust Visual SLAM
por: Du, Yanwei, et al.
Publicado: (2025) -
VISTA: Generative Visual Imagination for Vision-and-Language Navigation
por: Huang, Yanjia, et al.
Publicado: (2025) -
VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation
por: Zhang, Chaofan, et al.
Publicado: (2025) -
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
por: Glossop, Catherine, et al.
Publicado: (2025)