ET tu, CLIP? Addressing Common Object Errors for Unseen Environments
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910502009962496 |
|---|---|
| author | Byun, Ye Won Jiao, Cathy Noroozizadeh, Shahriar Sun, Jimin Vitiello, Rosa |
| author_facet | Byun, Ye Won Jiao, Cathy Noroozizadeh, Shahriar Sun, Jimin Vitiello, Rosa |
| contents | We introduce a simple method that employs pre-trained CLIP encoders to enhance model generalization in the ALFRED task. In contrast to previous literature where CLIP replaces the visual encoder, we suggest using CLIP as an additional module through an auxiliary object detection objective. We validate our method on the recently proposed Episodic Transformer architecture and demonstrate that incorporating CLIP improves task performance on the unseen validation set. Additionally, our analysis results support that CLIP especially helps with leveraging object descriptions, detecting small objects, and interpreting rare words. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_17876 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ET tu, CLIP? Addressing Common Object Errors for Unseen Environments Byun, Ye Won Jiao, Cathy Noroozizadeh, Shahriar Sun, Jimin Vitiello, Rosa Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning Robotics We introduce a simple method that employs pre-trained CLIP encoders to enhance model generalization in the ALFRED task. In contrast to previous literature where CLIP replaces the visual encoder, we suggest using CLIP as an additional module through an auxiliary object detection objective. We validate our method on the recently proposed Episodic Transformer architecture and demonstrate that incorporating CLIP improves task performance on the unseen validation set. Additionally, our analysis results support that CLIP especially helps with leveraging object descriptions, detecting small objects, and interpreting rare words. |
| title | ET tu, CLIP? Addressing Common Object Errors for Unseen Environments |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning Robotics |
| url | https://arxiv.org/abs/2406.17876 |