ET tu, CLIP? Addressing Common Object Errors for Unseen Environments
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910502009962496 |
|---|---|
| author | Byun, Ye Won Jiao, Cathy Noroozizadeh, Shahriar Sun, Jimin Vitiello, Rosa |
| author_facet | Byun, Ye Won Jiao, Cathy Noroozizadeh, Shahriar Sun, Jimin Vitiello, Rosa |
| contents | We introduce a simple method that employs pre-trained CLIP encoders to enhance model generalization in the ALFRED task. In contrast to previous literature where CLIP replaces the visual encoder, we suggest using CLIP as an additional module through an auxiliary object detection objective. We validate our method on the recently proposed Episodic Transformer architecture and demonstrate that incorporating CLIP improves task performance on the unseen validation set. Additionally, our analysis results support that CLIP especially helps with leveraging object descriptions, detecting small objects, and interpreting rare words. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_17876 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ET tu, CLIP? Addressing Common Object Errors for Unseen Environments Byun, Ye Won Jiao, Cathy Noroozizadeh, Shahriar Sun, Jimin Vitiello, Rosa Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning Robotics We introduce a simple method that employs pre-trained CLIP encoders to enhance model generalization in the ALFRED task. In contrast to previous literature where CLIP replaces the visual encoder, we suggest using CLIP as an additional module through an auxiliary object detection objective. We validate our method on the recently proposed Episodic Transformer architecture and demonstrate that incorporating CLIP improves task performance on the unseen validation set. Additionally, our analysis results support that CLIP especially helps with leveraging object descriptions, detecting small objects, and interpreting rare words. |
| title | ET tu, CLIP? Addressing Common Object Errors for Unseen Environments |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning Robotics |
| url | https://arxiv.org/abs/2406.17876 |