Embodied Domain Adaptation for Object Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shi, Xiangyu, Qiao, Yanyuan, Liu, Lingqiao, Dayoub, Feras
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916812536414208
author Shi, Xiangyu
Qiao, Yanyuan
Liu, Lingqiao
Dayoub, Feras
author_facet Shi, Xiangyu
Qiao, Yanyuan
Liu, Lingqiao
Dayoub, Feras
contents Mobile robots rely on object detectors for perception and object localization in indoor environments. However, standard closed-set methods struggle to handle the diverse objects and dynamic conditions encountered in real homes and labs. Open-vocabulary object detection (OVOD), driven by Vision Language Models (VLMs), extends beyond fixed labels but still struggles with domain shifts in indoor environments. We introduce a Source-Free Domain Adaptation (SFDA) approach that adapts a pre-trained model without accessing source data. We refine pseudo labels via temporal clustering, employ multi-scale threshold fusion, and apply a Mean Teacher framework with contrastive learning. Our Embodied Domain Adaptation for Object Detection (EDAOD) benchmark evaluates adaptation under sequential changes in lighting, layout, and object diversity. Our experiments show significant gains in zero-shot detection performance and flexible adaptation to dynamic indoor conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21860
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Embodied Domain Adaptation for Object Detection
Shi, Xiangyu
Qiao, Yanyuan
Liu, Lingqiao
Dayoub, Feras
Robotics
Computer Vision and Pattern Recognition
Mobile robots rely on object detectors for perception and object localization in indoor environments. However, standard closed-set methods struggle to handle the diverse objects and dynamic conditions encountered in real homes and labs. Open-vocabulary object detection (OVOD), driven by Vision Language Models (VLMs), extends beyond fixed labels but still struggles with domain shifts in indoor environments. We introduce a Source-Free Domain Adaptation (SFDA) approach that adapts a pre-trained model without accessing source data. We refine pseudo labels via temporal clustering, employ multi-scale threshold fusion, and apply a Mean Teacher framework with contrastive learning. Our Embodied Domain Adaptation for Object Detection (EDAOD) benchmark evaluates adaptation under sequential changes in lighting, layout, and object diversity. Our experiments show significant gains in zero-shot detection performance and flexible adaptation to dynamic indoor conditions.
title Embodied Domain Adaptation for Object Detection
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.21860