See and Switch: Vision-Based Branching for Interactive Robot-Skill Programming
Fuente:
arXiv
Saved in:
| Main Authors: | Vanc, Petr, Behrens, Jan Kristof, Hlaváč, Václav, Stepanova, Karla |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Should We Teach Robots? A Comparison of Kinesthetic, Joystick, and Gesture-Based Teaching
by: Vanc, Petr, et al.
Published: (2026)
by: Vanc, Petr, et al.
Published: (2026)
Communicating human intent to a robotic companion by multi-type gesture sentences
by: Vanc, Petr, et al.
Published: (2023)
by: Vanc, Petr, et al.
Published: (2023)
ILeSiA: Interactive Learning of Robot Situational Awareness from Camera Input
by: Vanc, Petr, et al.
Published: (2024)
by: Vanc, Petr, et al.
Published: (2024)
Closed Loop Interactive Embodied Reasoning for Robot Manipulation
by: Nazarczuk, Michal, et al.
Published: (2024)
by: Nazarczuk, Michal, et al.
Published: (2024)
Context-aware robot control using gesture episodes
by: Vanc, Petr, et al.
Published: (2023)
by: Vanc, Petr, et al.
Published: (2023)
TransforMerger: Transformer-based Voice-Gesture Fusion for Robust Human-Robot Communication
by: Vanc, Petr, et al.
Published: (2025)
by: Vanc, Petr, et al.
Published: (2025)
Tell and show: Combining multiple modalities to communicate manipulation tasks to a robot
by: Vanc, Petr, et al.
Published: (2024)
by: Vanc, Petr, et al.
Published: (2024)
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
by: Dai, Tingjun, et al.
Published: (2026)
by: Dai, Tingjun, et al.
Published: (2026)
Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
by: Kerr, Justin, et al.
Published: (2024)
by: Kerr, Justin, et al.
Published: (2024)
How Robot Dogs See the Unseeable: Improving Visual Interpretability via Peering for Exploratory Robots
by: Bimber, Oliver, et al.
Published: (2025)
by: Bimber, Oliver, et al.
Published: (2025)
AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents
by: Cui, Jieming, et al.
Published: (2024)
by: Cui, Jieming, et al.
Published: (2024)
Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
by: Bai, Yongjie, et al.
Published: (2025)
by: Bai, Yongjie, et al.
Published: (2025)
Vision-Based Safe Human-Robot Collaboration with Uncertainty Guarantees
by: Thumm, Jakob, et al.
Published: (2026)
by: Thumm, Jakob, et al.
Published: (2026)
Multi-Platform Teach-and-Repeat Navigation by Visual Place Recognition Based on Deep-Learned Local Features
by: Truhlařík, Václav, et al.
Published: (2025)
by: Truhlařík, Václav, et al.
Published: (2025)
HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction
by: Shi, Zhonghao, et al.
Published: (2025)
by: Shi, Zhonghao, et al.
Published: (2025)
Vision Language Action Models in Robotic Manipulation: A Systematic Review
by: Din, Muhayy Ud, et al.
Published: (2025)
by: Din, Muhayy Ud, et al.
Published: (2025)
VANP: Learning Where to See for Navigation with Self-Supervised Vision-Action Pre-Training
by: Nazeri, Mohammad, et al.
Published: (2024)
by: Nazeri, Mohammad, et al.
Published: (2024)
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
by: Xu, Kechun, et al.
Published: (2025)
by: Xu, Kechun, et al.
Published: (2025)
UniPrototype: Humn-Robot Skill Learning with Uniform Prototypes
by: Hu, Xiao, et al.
Published: (2025)
by: Hu, Xiao, et al.
Published: (2025)
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
by: Feng, Yixu, et al.
Published: (2026)
by: Feng, Yixu, et al.
Published: (2026)
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
by: Hui, Chenyu, et al.
Published: (2026)
by: Hui, Chenyu, et al.
Published: (2026)
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
by: Cheng, An-Chieh, et al.
Published: (2024)
by: Cheng, An-Chieh, et al.
Published: (2024)
Whom to Respond To? A Transformer-Based Model for Multi-Party Social Robot Interaction
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
by: Han, Xiaofeng, et al.
Published: (2025)
by: Han, Xiaofeng, et al.
Published: (2025)
HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
by: Koo, Myungkyu, et al.
Published: (2025)
by: Koo, Myungkyu, et al.
Published: (2025)
Towards an Accurate and Effective Robot Vision (The Problem of Topological Localization for Mobile Robots)
by: Boros, Emanuela
Published: (2025)
by: Boros, Emanuela
Published: (2025)
Enhancing Reusability of Learned Skills for Robot Manipulation via Gaze Information and Motion Bottlenecks
by: Takizawa, Ryo, et al.
Published: (2025)
by: Takizawa, Ryo, et al.
Published: (2025)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
by: Li, Rong, et al.
Published: (2024)
by: Li, Rong, et al.
Published: (2024)
Multimodal Anomaly Detection for Human-Robot Interaction
by: Ribeiro, Guilherme, et al.
Published: (2026)
by: Ribeiro, Guilherme, et al.
Published: (2026)
See Silhouettes in Motion with Neuromorphic Vision
by: Zhang, Pei, et al.
Published: (2026)
by: Zhang, Pei, et al.
Published: (2026)
Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation
by: Zhang, Xitie, et al.
Published: (2026)
by: Zhang, Xitie, et al.
Published: (2026)
Integrating Persian Lip Reading in Surena-V Humanoid Robot for Human-Robot Interaction
by: Abbasi, Ali Farshian, et al.
Published: (2025)
by: Abbasi, Ali Farshian, et al.
Published: (2025)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
RobotPan: A 360$^\circ$ Surround-View Robotic Vision System for Embodied Perception
by: Ma, Jiahao, et al.
Published: (2026)
by: Ma, Jiahao, et al.
Published: (2026)
Depth Jitter: Seeing through the Depth
by: Rahman, Md Sazidur, et al.
Published: (2025)
by: Rahman, Md Sazidur, et al.
Published: (2025)
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
by: Wang, Beichen, et al.
Published: (2024)
by: Wang, Beichen, et al.
Published: (2024)
High-Definition 5MP Stereo Vision Sensing for Robotics
by: Jiang, Leaf, et al.
Published: (2026)
by: Jiang, Leaf, et al.
Published: (2026)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
ViTA-Seg: Vision Transformer for Amodal Segmentation in Robotics
by: Caramia, Donato, et al.
Published: (2025)
by: Caramia, Donato, et al.
Published: (2025)
Similar Items
-
How Should We Teach Robots? A Comparison of Kinesthetic, Joystick, and Gesture-Based Teaching
by: Vanc, Petr, et al.
Published: (2026) -
Communicating human intent to a robotic companion by multi-type gesture sentences
by: Vanc, Petr, et al.
Published: (2023) -
ILeSiA: Interactive Learning of Robot Situational Awareness from Camera Input
by: Vanc, Petr, et al.
Published: (2024) -
Closed Loop Interactive Embodied Reasoning for Robot Manipulation
by: Nazarczuk, Michal, et al.
Published: (2024) -
Context-aware robot control using gesture episodes
by: Vanc, Petr, et al.
Published: (2023)