Learning Visual Hierarchies in Hyperbolic Space for Image Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915709134569472 |
|---|---|
| author | Wang, Ziwei Ramasinghe, Sameera Xu, Chenchen Monteil, Julien Bazzani, Loris Ajanthan, Thalaiyasingam |
| author_facet | Wang, Ziwei Ramasinghe, Sameera Xu, Chenchen Monteil, Julien Bazzani, Loris Ajanthan, Thalaiyasingam |
| contents | Structuring latent representations in a hierarchical manner enables models to learn patterns at multiple levels of abstraction. However, most prevalent image understanding models focus on visual similarity, and learning visual hierarchies is relatively unexplored. In this work, for the first time, we introduce a learning paradigm that can encode user-defined multi-level complex visual hierarchies in hyperbolic space without requiring explicit hierarchical labels. As a concrete example, first, we define a part-based image hierarchy using object-level annotations within and across images. Then, we introduce an approach to enforce the hierarchy using contrastive loss with pairwise entailment metrics. Finally, we discuss new evaluation metrics to effectively measure hierarchical image retrieval. Encoding these complex relationships ensures that the learned representations capture semantic and structural information that transcends mere visual similarity. Experiments in part-based image retrieval show significant improvements in hierarchical retrieval tasks, demonstrating the capability of our model in capturing visual hierarchies. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_17490 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Learning Visual Hierarchies in Hyperbolic Space for Image Retrieval Wang, Ziwei Ramasinghe, Sameera Xu, Chenchen Monteil, Julien Bazzani, Loris Ajanthan, Thalaiyasingam Computer Vision and Pattern Recognition Structuring latent representations in a hierarchical manner enables models to learn patterns at multiple levels of abstraction. However, most prevalent image understanding models focus on visual similarity, and learning visual hierarchies is relatively unexplored. In this work, for the first time, we introduce a learning paradigm that can encode user-defined multi-level complex visual hierarchies in hyperbolic space without requiring explicit hierarchical labels. As a concrete example, first, we define a part-based image hierarchy using object-level annotations within and across images. Then, we introduce an approach to enforce the hierarchy using contrastive loss with pairwise entailment metrics. Finally, we discuss new evaluation metrics to effectively measure hierarchical image retrieval. Encoding these complex relationships ensures that the learned representations capture semantic and structural information that transcends mere visual similarity. Experiments in part-based image retrieval show significant improvements in hierarchical retrieval tasks, demonstrating the capability of our model in capturing visual hierarchies. |
| title | Learning Visual Hierarchies in Hyperbolic Space for Image Retrieval |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2411.17490 |