OTAS: Open-vocabulary Token Alignment for Outdoor Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schwaiger, Simon, Thalhammer, Stefan, Wöber, Wilfried, Steinbauer-Wagner, Gerald
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908551182548992
author Schwaiger, Simon
Thalhammer, Stefan
Wöber, Wilfried
Steinbauer-Wagner, Gerald
author_facet Schwaiger, Simon
Thalhammer, Stefan
Wöber, Wilfried
Steinbauer-Wagner, Gerald
contents Understanding open-world semantics is critical for robotic planning and control, particularly in unstructured outdoor environments. Existing vision-language mapping approaches typically rely on object-centric segmentation priors, which often fail outdoors due to semantic ambiguities and indistinct class boundaries. We propose OTAS - an Open-vocabulary Token Alignment method for outdoor Segmentation. OTAS addresses the limitations of open-vocabulary segmentation models by extracting semantic structure directly from the output tokens of pre-trained vision models. By clustering semantically similar structures across single and multiple views and grounding them in language, OTAS reconstructs a geometrically consistent feature field that supports open-vocabulary segmentation queries. Our method operates in a zero-shot manner, without scene-specific fine-tuning, and achieves real-time performance of up to ~17 fps. On the Off-Road Freespace Detection dataset, OTAS yields a modest IoU improvement over fine-tuned and open-vocabulary 2D segmentation baselines. In 3D segmentation on TartanAir, it achieves up to a 151% relative IoU improvement compared to existing open-vocabulary mapping methods. Real-world reconstructions further demonstrate OTAS' applicability to robotic deployment. Code and a ROS 2 node are available at https://otas-segmentation.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08851
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OTAS: Open-vocabulary Token Alignment for Outdoor Segmentation
Schwaiger, Simon
Thalhammer, Stefan
Wöber, Wilfried
Steinbauer-Wagner, Gerald
Robotics
Understanding open-world semantics is critical for robotic planning and control, particularly in unstructured outdoor environments. Existing vision-language mapping approaches typically rely on object-centric segmentation priors, which often fail outdoors due to semantic ambiguities and indistinct class boundaries. We propose OTAS - an Open-vocabulary Token Alignment method for outdoor Segmentation. OTAS addresses the limitations of open-vocabulary segmentation models by extracting semantic structure directly from the output tokens of pre-trained vision models. By clustering semantically similar structures across single and multiple views and grounding them in language, OTAS reconstructs a geometrically consistent feature field that supports open-vocabulary segmentation queries. Our method operates in a zero-shot manner, without scene-specific fine-tuning, and achieves real-time performance of up to ~17 fps. On the Off-Road Freespace Detection dataset, OTAS yields a modest IoU improvement over fine-tuned and open-vocabulary 2D segmentation baselines. In 3D segmentation on TartanAir, it achieves up to a 151% relative IoU improvement compared to existing open-vocabulary mapping methods. Real-world reconstructions further demonstrate OTAS' applicability to robotic deployment. Code and a ROS 2 node are available at https://otas-segmentation.github.io/.
title OTAS: Open-vocabulary Token Alignment for Outdoor Segmentation
topic Robotics
url https://arxiv.org/abs/2507.08851