Global and Local Entailment Learning for Natural World Imagery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sastry, Srikumar, Dhakal, Aayush, Xing, Eric, Khanal, Subash, Jacobs, Nathan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912451834937344
author Sastry, Srikumar
Dhakal, Aayush
Xing, Eric
Khanal, Subash
Jacobs, Nathan
author_facet Sastry, Srikumar
Dhakal, Aayush
Xing, Eric
Khanal, Subash
Jacobs, Nathan
contents Learning the hierarchical structure of data in vision-language models is a significant challenge. Previous works have attempted to address this challenge by employing entailment learning. However, these approaches fail to model the transitive nature of entailment explicitly, which establishes the relationship between order and semantics within a representation space. In this work, we introduce Radial Cross-Modal Embeddings (RCME), a framework that enables the explicit modeling of transitivity-enforced entailment. Our proposed framework optimizes for the partial order of concepts within vision-language models. By leveraging our framework, we develop a hierarchical vision-language foundation model capable of representing the hierarchy in the Tree of Life. Our experiments on hierarchical species classification and hierarchical retrieval tasks demonstrate the enhanced performance of our models compared to the existing state-of-the-art models. Our code and models are open-sourced at https://vishu26.github.io/RCME/index.html.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21476
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Global and Local Entailment Learning for Natural World Imagery
Sastry, Srikumar
Dhakal, Aayush
Xing, Eric
Khanal, Subash
Jacobs, Nathan
Computer Vision and Pattern Recognition
Learning the hierarchical structure of data in vision-language models is a significant challenge. Previous works have attempted to address this challenge by employing entailment learning. However, these approaches fail to model the transitive nature of entailment explicitly, which establishes the relationship between order and semantics within a representation space. In this work, we introduce Radial Cross-Modal Embeddings (RCME), a framework that enables the explicit modeling of transitivity-enforced entailment. Our proposed framework optimizes for the partial order of concepts within vision-language models. By leveraging our framework, we develop a hierarchical vision-language foundation model capable of representing the hierarchy in the Tree of Life. Our experiments on hierarchical species classification and hierarchical retrieval tasks demonstrate the enhanced performance of our models compared to the existing state-of-the-art models. Our code and models are open-sourced at https://vishu26.github.io/RCME/index.html.
title Global and Local Entailment Learning for Natural World Imagery
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.21476