Enabling Dynamic Tracking in Vision-Language-Action Models via Time-Discrete and Time-Continuous Velocity Feedforward
Fuente:
arXiv
Guardado en:
| Autores principales: | Hechtl, Johannes, Schmitt, Philipp, von Wichert, Georg, Burgard, Wolfram |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Factory-Floor Deployment Case Study of VLA Pipelines for Industrial Packaging Task: Workflow, Failures, and Lessons
por: Zhu, Brian, et al.
Publicado: (2026)
por: Zhu, Brian, et al.
Publicado: (2026)
Robot Control Stack: A Lean Ecosystem for Robot Learning at Scale
por: Jülg, Tobias, et al.
Publicado: (2025)
por: Jülg, Tobias, et al.
Publicado: (2025)
Rewarding DINO: Predicting Dense Rewards with Vision Foundation Models
por: Krack, Pierre, et al.
Publicado: (2026)
por: Krack, Pierre, et al.
Publicado: (2026)
Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models
por: Liu, Haoyun, et al.
Publicado: (2026)
por: Liu, Haoyun, et al.
Publicado: (2026)
VLM-Vac: Enhancing Smart Vacuums through VLM Knowledge Distillation and Language-Guided Experience Replay
por: Mirjalili, Reihaneh, et al.
Publicado: (2024)
por: Mirjalili, Reihaneh, et al.
Publicado: (2024)
uPLAM: Robust Panoptic Localization and Mapping Leveraging Perception Uncertainties
por: Sirohi, Kshitij, et al.
Publicado: (2024)
por: Sirohi, Kshitij, et al.
Publicado: (2024)
CloudTrack: Scalable UAV Tracking with Cloud Semantics
por: Blei, Yannik, et al.
Publicado: (2024)
por: Blei, Yannik, et al.
Publicado: (2024)
Learning Continuous Control with Geometric Regularity from Robot Intrinsic Symmetry
por: Yan, Shengchao, et al.
Publicado: (2023)
por: Yan, Shengchao, et al.
Publicado: (2023)
Vision-Based Autonomous UAV Navigation and Landing for Urban Search and Rescue
por: Mittal, Mayank, et al.
Publicado: (2019)
por: Mittal, Mayank, et al.
Publicado: (2019)
CoDEPS: Online Continual Learning for Depth Estimation and Panoptic Segmentation
por: Vödisch, Niclas, et al.
Publicado: (2023)
por: Vödisch, Niclas, et al.
Publicado: (2023)
Collaborative Dynamic 3D Scene Graphs for Automated Driving
por: Greve, Elias, et al.
Publicado: (2023)
por: Greve, Elias, et al.
Publicado: (2023)
Agent-Agnostic Centralized Training for Decentralized Multi-Agent Cooperative Driving
por: Yan, Shengchao, et al.
Publicado: (2024)
por: Yan, Shengchao, et al.
Publicado: (2024)
Refined Policy Distillation: From VLA Generalists to RL Experts
por: Jülg, Tobias, et al.
Publicado: (2025)
por: Jülg, Tobias, et al.
Publicado: (2025)
Lan-grasp: Using Large Language Models for Semantic Object Grasping and Placement
por: Mirjalili, Reihaneh, et al.
Publicado: (2023)
por: Mirjalili, Reihaneh, et al.
Publicado: (2023)
BYE: Build Your Encoder with One Sequence of Exploration Data for Long-Term Dynamic Scene Understanding
por: Huang, Chenguang, et al.
Publicado: (2024)
por: Huang, Chenguang, et al.
Publicado: (2024)
LiDAR Registration with Visual Foundation Models
por: Vödisch, Niclas, et al.
Publicado: (2025)
por: Vödisch, Niclas, et al.
Publicado: (2025)
ZAPP! Zonotope Agreement of Prediction and Planning for Continuous-Time Collision Avoidance with Discrete-Time Dynamics
por: Paparusso, Luca, et al.
Publicado: (2024)
por: Paparusso, Luca, et al.
Publicado: (2024)
Augmented Reality for RObots (ARRO): Pointing Visuomotor Policies Towards Visual Robustness
por: Mirjalili, Reihaneh, et al.
Publicado: (2025)
por: Mirjalili, Reihaneh, et al.
Publicado: (2025)
Automatic Target-Less Camera-LiDAR Calibration From Motion and Deep Point Correspondences
por: Petek, Kürsat, et al.
Publicado: (2024)
por: Petek, Kürsat, et al.
Publicado: (2024)
UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models
por: Zhang, Qiyao, et al.
Publicado: (2026)
por: Zhang, Qiyao, et al.
Publicado: (2026)
Concept-Based Dictionary Learning for Inference-Time Safety in Vision Language Action Models
por: Wen, Siqi, et al.
Publicado: (2026)
por: Wen, Siqi, et al.
Publicado: (2026)
RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action Model
por: Dai, Mingtong, et al.
Publicado: (2025)
por: Dai, Mingtong, et al.
Publicado: (2025)
ECHO: Continuous Hierarchical Memory for Vision-Language-Action Models
por: Hu, Yanbin, et al.
Publicado: (2026)
por: Hu, Yanbin, et al.
Publicado: (2026)
HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing
por: Gubernatorov, Konstantin, et al.
Publicado: (2026)
por: Gubernatorov, Konstantin, et al.
Publicado: (2026)
Benchmarking Empirical and Learning-Based Approaches for Feedforward Steering Control in Autonomous Racing
por: Jank, Georg, et al.
Publicado: (2026)
por: Jank, Georg, et al.
Publicado: (2026)
Verifier-free Test-Time Sampling for Vision Language Action Models
por: Jang, Suhyeok, et al.
Publicado: (2025)
por: Jang, Suhyeok, et al.
Publicado: (2025)
Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning
por: Tur, Yalcin, et al.
Publicado: (2026)
por: Tur, Yalcin, et al.
Publicado: (2026)
Bayesian Optimization for Sample-Efficient Policy Improvement in Robotic Manipulation
por: Röfer, Adrian, et al.
Publicado: (2024)
por: Röfer, Adrian, et al.
Publicado: (2024)
Bearing-Only Tracking and Circumnavigation of a Fast Time-Varied Velocity Target Utilising an LSTM
por: Torok, Mitchell, et al.
Publicado: (2025)
por: Torok, Mitchell, et al.
Publicado: (2025)
Continually Evolving Skill Knowledge in Vision Language Action Model
por: Wu, Yuxuan, et al.
Publicado: (2025)
por: Wu, Yuxuan, et al.
Publicado: (2025)
Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
por: Li, Pengxiang, et al.
Publicado: (2025)
por: Li, Pengxiang, et al.
Publicado: (2025)
Few-Shot Panoptic Segmentation With Foundation Models
por: Käppeler, Markus, et al.
Publicado: (2023)
por: Käppeler, Markus, et al.
Publicado: (2023)
ConPoSe: LLM-Guided Contact Point Selection for Scalable Cooperative Object Pushing
por: Steinkrüger, Noah, et al.
Publicado: (2025)
por: Steinkrüger, Noah, et al.
Publicado: (2025)
Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
por: Cai, Rui, et al.
Publicado: (2026)
por: Cai, Rui, et al.
Publicado: (2026)
Test-Time Training for Visual Foresight Vision-Language-Action Models
por: Park, Sangwu, et al.
Publicado: (2026)
por: Park, Sangwu, et al.
Publicado: (2026)
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
por: Yang, Siyuan, et al.
Publicado: (2025)
por: Yang, Siyuan, et al.
Publicado: (2025)
Continuous Reasoning for Vision-Language-Action
por: Wu, Yueh-Hua, et al.
Publicado: (2026)
por: Wu, Yueh-Hua, et al.
Publicado: (2026)
Metamorphic Testing of Vision-Language Action-Enabled Robots
por: Valle, Pablo, et al.
Publicado: (2026)
por: Valle, Pablo, et al.
Publicado: (2026)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
por: Chen, Jiayi, et al.
Publicado: (2025)
por: Chen, Jiayi, et al.
Publicado: (2025)
Multimodal Spatial Language Maps for Robot Navigation and Manipulation
por: Huang, Chenguang, et al.
Publicado: (2025)
por: Huang, Chenguang, et al.
Publicado: (2025)
Ejemplares similares
-
A Factory-Floor Deployment Case Study of VLA Pipelines for Industrial Packaging Task: Workflow, Failures, and Lessons
por: Zhu, Brian, et al.
Publicado: (2026) -
Robot Control Stack: A Lean Ecosystem for Robot Learning at Scale
por: Jülg, Tobias, et al.
Publicado: (2025) -
Rewarding DINO: Predicting Dense Rewards with Vision Foundation Models
por: Krack, Pierre, et al.
Publicado: (2026) -
Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models
por: Liu, Haoyun, et al.
Publicado: (2026) -
VLM-Vac: Enhancing Smart Vacuums through VLM Knowledge Distillation and Language-Guided Experience Replay
por: Mirjalili, Reihaneh, et al.
Publicado: (2024)