10 Open Challenges Steering the Future of Vision-Language-Action Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Poria, Soujanya, Majumder, Navonil, Hung, Chia-Yu, Bagherzadeh, Amir Ali, Li, Chuan, Kwok, Kenneth, Wang, Ziwei, Tan, Cheston, Wu, Jiajun, Hsu, David |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
par: Hung, Chia-Yu, et autres
Publié: (2025)
par: Hung, Chia-Yu, et autres
Publié: (2025)
NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards
par: Hung, Chia-Yu, et autres
Publié: (2025)
par: Hung, Chia-Yu, et autres
Publié: (2025)
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
par: Liu, Renhang, et autres
Publié: (2025)
par: Liu, Renhang, et autres
Publié: (2025)
Inference Time Alignment with Reward-Guided Tree Search
par: Hung, Chia-Yu, et autres
Publié: (2024)
par: Hung, Chia-Yu, et autres
Publié: (2024)
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
par: Hung, Chia-Yu, et autres
Publié: (2024)
par: Hung, Chia-Yu, et autres
Publié: (2024)
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
par: Majumder, Navonil, et autres
Publié: (2024)
par: Majumder, Navonil, et autres
Publié: (2024)
From Grounding to Manipulation: Case Studies of Foundation Model Integration in Embodied Robotic Systems
par: Sui, Xiuchao, et autres
Publié: (2025)
par: Sui, Xiuchao, et autres
Publié: (2025)
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
par: Chia, Yew Ken, et autres
Publié: (2024)
par: Chia, Yew Ken, et autres
Publié: (2024)
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
par: Lin, Zijun, et autres
Publié: (2025)
par: Lin, Zijun, et autres
Publié: (2025)
SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
par: Gao, Tian, et autres
Publié: (2026)
par: Gao, Tian, et autres
Publié: (2026)
Mustango: Toward Controllable Text-to-Music Generation
par: Melechovsky, Jan, et autres
Publié: (2023)
par: Melechovsky, Jan, et autres
Publié: (2023)
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
par: Hong, Pengfei, et autres
Publié: (2024)
par: Hong, Pengfei, et autres
Publié: (2024)
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
par: Sun, Qi, et autres
Publié: (2024)
par: Sun, Qi, et autres
Publié: (2024)
Lessons from Training Grounded LLMs with Verifiable Rewards
par: Sim, Shang Hong, et autres
Publié: (2025)
par: Sim, Shang Hong, et autres
Publié: (2025)
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
par: Song, Maojia, et autres
Publié: (2024)
par: Song, Maojia, et autres
Publié: (2024)
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
par: Hu, Tianshuai, et autres
Publié: (2025)
par: Hu, Tianshuai, et autres
Publié: (2025)
Dexbotic: Open-Source Vision-Language-Action Toolbox
par: Xie, Bin, et autres
Publié: (2025)
par: Xie, Bin, et autres
Publié: (2025)
OpenVLA: An Open-Source Vision-Language-Action Model
par: Kim, Moo Jin, et autres
Publié: (2024)
par: Kim, Moo Jin, et autres
Publié: (2024)
Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms
par: Li, Qi, et autres
Publié: (2026)
par: Li, Qi, et autres
Publié: (2026)
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
par: Xu, Xiaoxu, et autres
Publié: (2026)
par: Xu, Xiaoxu, et autres
Publié: (2026)
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
par: Yang, Siyuan, et autres
Publié: (2025)
par: Yang, Siyuan, et autres
Publié: (2025)
Action-Constrained Imitation Learning
par: Yeh, Chia-Han, et autres
Publié: (2025)
par: Yeh, Chia-Han, et autres
Publié: (2025)
An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
par: Xu, Chao, et autres
Publié: (2025)
par: Xu, Chao, et autres
Publié: (2025)
A Survey on Path Planning Problem of Rolling Contacts: Approaches, Applications and Future Challenges
par: Tafrishi, Seyed Amir, et autres
Publié: (2025)
par: Tafrishi, Seyed Amir, et autres
Publié: (2025)
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
par: Zhao, Wei, et autres
Publié: (2025)
par: Zhao, Wei, et autres
Publié: (2025)
VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
par: Guo, Wenkai, et autres
Publié: (2025)
par: Guo, Wenkai, et autres
Publié: (2025)
Improving Text-To-Audio Models with Synthetic Captions
par: Kong, Zhifeng, et autres
Publié: (2024)
par: Kong, Zhifeng, et autres
Publié: (2024)
Do What You Say: Steering Vision-Language-Action Models via Runtime Reasoning-Action Alignment Verification
par: Wu, Yilin, et autres
Publié: (2025)
par: Wu, Yilin, et autres
Publié: (2025)
TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
par: Zhang, Zongzheng, et autres
Publié: (2025)
par: Zhang, Zongzheng, et autres
Publié: (2025)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
par: Li, Runhao, et autres
Publié: (2025)
par: Li, Runhao, et autres
Publié: (2025)
ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models
par: Li, Ye, et autres
Publié: (2026)
par: Li, Ye, et autres
Publié: (2026)
VICtoR: Learning Hierarchical Vision-Instruction Correlation Rewards for Long-horizon Manipulation
par: Hung, Kuo-Han, et autres
Publié: (2024)
par: Hung, Kuo-Han, et autres
Publié: (2024)
TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation
par: Huang, Yuzhe, et autres
Publié: (2026)
par: Huang, Yuzhe, et autres
Publié: (2026)
RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI
par: Tai, Cong, et autres
Publié: (2025)
par: Tai, Cong, et autres
Publié: (2025)
Contrastive Conceptor Activation Steering (COAST): Unlocking Vision-Language-Action Models through Hidden States
par: Miao, Miranda Muqing, et autres
Publié: (2026)
par: Miao, Miranda Muqing, et autres
Publié: (2026)
Steering Flexible Linear Objects in Planar Environments by Two Robot Hands Using Euler's Elastica Solutions
par: Levin, Aharon, et autres
Publié: (2025)
par: Levin, Aharon, et autres
Publié: (2025)
Open Scene Graphs for Open World Object-Goal Navigation
par: Loo, Joel, et autres
Publié: (2024)
par: Loo, Joel, et autres
Publié: (2024)
Open Scene Graphs for Open-World Object-Goal Navigation
par: Loo, Joel, et autres
Publié: (2025)
par: Loo, Joel, et autres
Publié: (2025)
Metamorphic Testing of Vision-Language Action-Enabled Robots
par: Valle, Pablo, et autres
Publié: (2026)
par: Valle, Pablo, et autres
Publié: (2026)
Scene Action Maps: Behavioural Maps for Navigation without Metric Information
par: Loo, Joel, et autres
Publié: (2024)
par: Loo, Joel, et autres
Publié: (2024)
Documents similaires
-
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
par: Hung, Chia-Yu, et autres
Publié: (2025) -
NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards
par: Hung, Chia-Yu, et autres
Publié: (2025) -
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
par: Liu, Renhang, et autres
Publié: (2025) -
Inference Time Alignment with Reward-Guided Tree Search
par: Hung, Chia-Yu, et autres
Publié: (2024) -
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
par: Hung, Chia-Yu, et autres
Publié: (2024)