ViT-VS: On the Applicability of Pretrained Vision Transformer Features for Generalizable Visual Servoing
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911051165990912 |
|---|---|
| author | Scherl, Alessandro Thalhammer, Stefan Neuberger, Bernhard Wöber, Wilfried García-Rodríguez, José |
| author_facet | Scherl, Alessandro Thalhammer, Stefan Neuberger, Bernhard Wöber, Wilfried García-Rodríguez, José |
| contents | Visual servoing enables robots to precisely position their end-effector relative to a target object. While classical methods rely on hand-crafted features and thus are universally applicable without task-specific training, they often struggle with occlusions and environmental variations, whereas learning-based approaches improve robustness but typically require extensive training. We present a visual servoing approach that leverages pretrained vision transformers for semantic feature extraction, combining the advantages of both paradigms while also being able to generalize beyond the provided sample. Our approach achieves full convergence in unperturbed scenarios and surpasses classical image-based visual servoing by up to 31.2\% relative improvement in perturbed scenarios. Even the convergence rates of learning-based methods are matched despite requiring no task- or object-specific training. Real-world evaluations confirm robust performance in end-effector positioning, industrial box manipulation, and grasping of unseen objects using only a reference from the same category. Our code and simulation environment are available at: https://alessandroscherl.github.io/ViT-VS/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_04545 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ViT-VS: On the Applicability of Pretrained Vision Transformer Features for Generalizable Visual Servoing Scherl, Alessandro Thalhammer, Stefan Neuberger, Bernhard Wöber, Wilfried García-Rodríguez, José Robotics Computer Vision and Pattern Recognition Visual servoing enables robots to precisely position their end-effector relative to a target object. While classical methods rely on hand-crafted features and thus are universally applicable without task-specific training, they often struggle with occlusions and environmental variations, whereas learning-based approaches improve robustness but typically require extensive training. We present a visual servoing approach that leverages pretrained vision transformers for semantic feature extraction, combining the advantages of both paradigms while also being able to generalize beyond the provided sample. Our approach achieves full convergence in unperturbed scenarios and surpasses classical image-based visual servoing by up to 31.2\% relative improvement in perturbed scenarios. Even the convergence rates of learning-based methods are matched despite requiring no task- or object-specific training. Real-world evaluations confirm robust performance in end-effector positioning, industrial box manipulation, and grasping of unseen objects using only a reference from the same category. Our code and simulation environment are available at: https://alessandroscherl.github.io/ViT-VS/ |
| title | ViT-VS: On the Applicability of Pretrained Vision Transformer Features for Generalizable Visual Servoing |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.04545 |