ET tu, CLIP? Addressing Common Object Errors for Unseen Environments

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Byun, Ye Won, Jiao, Cathy, Noroozizadeh, Shahriar, Sun, Jimin, Vitiello, Rosa
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910502009962496
author Byun, Ye Won
Jiao, Cathy
Noroozizadeh, Shahriar
Sun, Jimin
Vitiello, Rosa
author_facet Byun, Ye Won
Jiao, Cathy
Noroozizadeh, Shahriar
Sun, Jimin
Vitiello, Rosa
contents We introduce a simple method that employs pre-trained CLIP encoders to enhance model generalization in the ALFRED task. In contrast to previous literature where CLIP replaces the visual encoder, we suggest using CLIP as an additional module through an auxiliary object detection objective. We validate our method on the recently proposed Episodic Transformer architecture and demonstrate that incorporating CLIP improves task performance on the unseen validation set. Additionally, our analysis results support that CLIP especially helps with leveraging object descriptions, detecting small objects, and interpreting rare words.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17876
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ET tu, CLIP? Addressing Common Object Errors for Unseen Environments
Byun, Ye Won
Jiao, Cathy
Noroozizadeh, Shahriar
Sun, Jimin
Vitiello, Rosa
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Robotics
We introduce a simple method that employs pre-trained CLIP encoders to enhance model generalization in the ALFRED task. In contrast to previous literature where CLIP replaces the visual encoder, we suggest using CLIP as an additional module through an auxiliary object detection objective. We validate our method on the recently proposed Episodic Transformer architecture and demonstrate that incorporating CLIP improves task performance on the unseen validation set. Additionally, our analysis results support that CLIP especially helps with leveraging object descriptions, detecting small objects, and interpreting rare words.
title ET tu, CLIP? Addressing Common Object Errors for Unseen Environments
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Robotics
url https://arxiv.org/abs/2406.17876