Jointly Modeling Inter- & Intra-Modality Dependencies for Multi-modal Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913599345131520 |
|---|---|
| author | Madaan, Divyam Makino, Taro Chopra, Sumit Cho, Kyunghyun |
| author_facet | Madaan, Divyam Makino, Taro Chopra, Sumit Cho, Kyunghyun |
| contents | Supervised multi-modal learning involves mapping multiple modalities to a target label. Previous studies in this field have concentrated on capturing in isolation either the inter-modality dependencies (the relationships between different modalities and the label) or the intra-modality dependencies (the relationships within a single modality and the label). We argue that these conventional approaches that rely solely on either inter- or intra-modality dependencies may not be optimal in general. We view the multi-modal learning problem from the lens of generative models where we consider the target as a source of multiple modalities and the interaction between them. Towards that end, we propose inter- & intra-modality modeling (I2M2) framework, which captures and integrates both the inter- and intra-modality dependencies, leading to more accurate predictions. We evaluate our approach using real-world healthcare and vision-and-language datasets with state-of-the-art models, demonstrating superior performance over traditional methods focusing only on one type of modality dependency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_17613 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Jointly Modeling Inter- & Intra-Modality Dependencies for Multi-modal Learning Madaan, Divyam Makino, Taro Chopra, Sumit Cho, Kyunghyun Computer Vision and Pattern Recognition Computation and Language Machine Learning Supervised multi-modal learning involves mapping multiple modalities to a target label. Previous studies in this field have concentrated on capturing in isolation either the inter-modality dependencies (the relationships between different modalities and the label) or the intra-modality dependencies (the relationships within a single modality and the label). We argue that these conventional approaches that rely solely on either inter- or intra-modality dependencies may not be optimal in general. We view the multi-modal learning problem from the lens of generative models where we consider the target as a source of multiple modalities and the interaction between them. Towards that end, we propose inter- & intra-modality modeling (I2M2) framework, which captures and integrates both the inter- and intra-modality dependencies, leading to more accurate predictions. We evaluate our approach using real-world healthcare and vision-and-language datasets with state-of-the-art models, demonstrating superior performance over traditional methods focusing only on one type of modality dependency. |
| title | Jointly Modeling Inter- & Intra-Modality Dependencies for Multi-modal Learning |
| topic | Computer Vision and Pattern Recognition Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2405.17613 |