Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Wei, Fan, Xiaomeng, Wu, Yuwei, Gao, Zhi, Li, Pengxiang, Jia, Yunde, Harandi, Mehrtash
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913170267832320
author Wu, Wei
Fan, Xiaomeng
Wu, Yuwei
Gao, Zhi
Li, Pengxiang
Jia, Yunde
Harandi, Mehrtash
author_facet Wu, Wei
Fan, Xiaomeng
Wu, Yuwei
Gao, Zhi
Li, Pengxiang
Jia, Yunde
Harandi, Mehrtash
contents Modality alignment is critical for vision-language models (VLMs) to effectively integrate information across modalities. However, existing methods extract hierarchical features from text while representing each image with a single feature, leading to asymmetric and suboptimal alignment. To address this, we propose Alignment across Trees, a method that constructs and aligns tree-like hierarchical features for both image and text modalities. Specifically, we introduce a semantic-aware visual feature extraction framework that applies a cross-attention mechanism to visual class tokens from intermediate Transformer layers, guided by textual cues to extract visual features with coarse-to-fine semantics. We then embed the feature trees of the two modalities into hyperbolic manifolds with distinct curvatures to effectively model their hierarchical structures. To align across the heterogeneous hyperbolic manifolds with different curvatures, we formulate a KL distance measure between distributions on heterogeneous manifolds, and learn an intermediate manifold for manifold alignment by minimizing the distance. We prove the existence and uniqueness of the optimal intermediate manifold. Experiments on taxonomic open-set classification tasks across multiple image datasets demonstrate that our method consistently outperforms strong baselines under few-shot and cross-domain settings.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27391
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds
Wu, Wei
Fan, Xiaomeng
Wu, Yuwei
Gao, Zhi
Li, Pengxiang
Jia, Yunde
Harandi, Mehrtash
Computer Vision and Pattern Recognition
Machine Learning
Modality alignment is critical for vision-language models (VLMs) to effectively integrate information across modalities. However, existing methods extract hierarchical features from text while representing each image with a single feature, leading to asymmetric and suboptimal alignment. To address this, we propose Alignment across Trees, a method that constructs and aligns tree-like hierarchical features for both image and text modalities. Specifically, we introduce a semantic-aware visual feature extraction framework that applies a cross-attention mechanism to visual class tokens from intermediate Transformer layers, guided by textual cues to extract visual features with coarse-to-fine semantics. We then embed the feature trees of the two modalities into hyperbolic manifolds with distinct curvatures to effectively model their hierarchical structures. To align across the heterogeneous hyperbolic manifolds with different curvatures, we formulate a KL distance measure between distributions on heterogeneous manifolds, and learn an intermediate manifold for manifold alignment by minimizing the distance. We prove the existence and uniqueness of the optimal intermediate manifold. Experiments on taxonomic open-set classification tasks across multiple image datasets demonstrate that our method consistently outperforms strong baselines under few-shot and cross-domain settings.
title Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2510.27391