Show, Don't Tell: Detecting Novel Objects by Watching Human Videos

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Akl, James, Arbelaez, Jose Nicolas Avendano, Barabas, James, Barry, Jennifer L., Ching, Kalie, Eshed, Noam, Fu, Jiahui, Hidalgo, Michel, Hoelscher, Andrew, Kusnur, Tushar, Messing, Andrew, Nagler, Zachary, Okorn, Brian, Passerino, Mauro, Perkins, Tim J., Rosen, Eric, Shah, Ankit, Shankar, Tanmay, Shaw, Scott
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918386607325184
author Akl, James
Arbelaez, Jose Nicolas Avendano
Barabas, James
Barry, Jennifer L.
Ching, Kalie
Eshed, Noam
Fu, Jiahui
Hidalgo, Michel
Hoelscher, Andrew
Kusnur, Tushar
Messing, Andrew
Nagler, Zachary
Okorn, Brian
Passerino, Mauro
Perkins, Tim J.
Rosen, Eric
Shah, Ankit
Shankar, Tanmay
Shaw, Scott
author_facet Akl, James
Arbelaez, Jose Nicolas Avendano
Barabas, James
Barry, Jennifer L.
Ching, Kalie
Eshed, Noam
Fu, Jiahui
Hidalgo, Michel
Hoelscher, Andrew
Kusnur, Tushar
Messing, Andrew
Nagler, Zachary
Okorn, Brian
Passerino, Mauro
Perkins, Tim J.
Rosen, Eric
Shah, Ankit
Shankar, Tanmay
Shaw, Scott
contents How can a robot quickly identify and recognize new objects shown to it during a human demonstration? Existing closed-set object detectors frequently fail at this because the objects are out-of-distribution. While open-set detectors (e.g., VLMs) sometimes succeed, they often require expensive and tedious human-in-the-loop prompt engineering to uniquely recognize novel object instances. In this paper, we present a self-supervised system that eliminates the need for tedious language descriptions and expensive prompt engineering by training a bespoke object detector on an automatically created dataset, supervised by the human demonstration itself. In our approach, "Show, Don't Tell," we show the detector the specific objects of interest during the demonstration, rather than telling the detector about these objects via complex language descriptions. By bypassing language altogether, this paradigm enables us to quickly train bespoke detectors tailored to the relevant objects observed in human task demonstrations. We develop an integrated on-robot system to deploy our "Show, Don't Tell" paradigm of automatic dataset creation and novel object-detection on a real-world robot. Empirical results demonstrate that our pipeline significantly outperforms state-of-the-art detection and recognition methods for manipulated objects, leading to improved task completion for the robot.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12751
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Show, Don't Tell: Detecting Novel Objects by Watching Human Videos
Akl, James
Arbelaez, Jose Nicolas Avendano
Barabas, James
Barry, Jennifer L.
Ching, Kalie
Eshed, Noam
Fu, Jiahui
Hidalgo, Michel
Hoelscher, Andrew
Kusnur, Tushar
Messing, Andrew
Nagler, Zachary
Okorn, Brian
Passerino, Mauro
Perkins, Tim J.
Rosen, Eric
Shah, Ankit
Shankar, Tanmay
Shaw, Scott
Computer Vision and Pattern Recognition
Machine Learning
Robotics
How can a robot quickly identify and recognize new objects shown to it during a human demonstration? Existing closed-set object detectors frequently fail at this because the objects are out-of-distribution. While open-set detectors (e.g., VLMs) sometimes succeed, they often require expensive and tedious human-in-the-loop prompt engineering to uniquely recognize novel object instances. In this paper, we present a self-supervised system that eliminates the need for tedious language descriptions and expensive prompt engineering by training a bespoke object detector on an automatically created dataset, supervised by the human demonstration itself. In our approach, "Show, Don't Tell," we show the detector the specific objects of interest during the demonstration, rather than telling the detector about these objects via complex language descriptions. By bypassing language altogether, this paradigm enables us to quickly train bespoke detectors tailored to the relevant objects observed in human task demonstrations. We develop an integrated on-robot system to deploy our "Show, Don't Tell" paradigm of automatic dataset creation and novel object-detection on a real-world robot. Empirical results demonstrate that our pipeline significantly outperforms state-of-the-art detection and recognition methods for manipulated objects, leading to improved task completion for the robot.
title Show, Don't Tell: Detecting Novel Objects by Watching Human Videos
topic Computer Vision and Pattern Recognition
Machine Learning
Robotics
url https://arxiv.org/abs/2603.12751