Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Jiale, Zhou, Haoming, Liu, Yishu, Chen, Bingzhi, Jiang, Yuncheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification
von: Liu, Yangyang, et al.
Veröffentlicht: (2025)
von: Liu, Yangyang, et al.
Veröffentlicht: (2025)
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
MSCT: Differential Cross-Modal Attention for Deepfake Detection
von: Wei, Fangda, et al.
Veröffentlicht: (2026)
von: Wei, Fangda, et al.
Veröffentlicht: (2026)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
von: Rinaldi, Ivan, et al.
Veröffentlicht: (2026)
von: Rinaldi, Ivan, et al.
Veröffentlicht: (2026)
CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer
von: Wang, Yabing, et al.
Veröffentlicht: (2023)
von: Wang, Yabing, et al.
Veröffentlicht: (2023)
Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching
von: Zhang, Yafei, et al.
Veröffentlicht: (2025)
von: Zhang, Yafei, et al.
Veröffentlicht: (2025)
DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval Guidelines
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
Learning Contrastive Self-Distillation for Ultra-Fine-Grained Visual Categorization Targeting Limited Samples
von: Fang, Ziye, et al.
Veröffentlicht: (2023)
von: Fang, Ziye, et al.
Veröffentlicht: (2023)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition
von: Huang, Shunyu, et al.
Veröffentlicht: (2026)
von: Huang, Shunyu, et al.
Veröffentlicht: (2026)
Wavelet-Decoupling Contrastive Enhancement Network for Fine-Grained Skeleton-Based Action Recognition
von: Chang, Haochen, et al.
Veröffentlicht: (2024)
von: Chang, Haochen, et al.
Veröffentlicht: (2024)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
Hypergraph Tversky-Aware Domain Incremental Learning for Brain Tumor Segmentation with Missing Modalities
von: Wang, Junze, et al.
Veröffentlicht: (2025)
von: Wang, Junze, et al.
Veröffentlicht: (2025)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
BiTDiff: Fine-Grained 3D Conducting Motion Generation via BiMamba-Transformer Diffusion
von: Jia, Tianzhi, et al.
Veröffentlicht: (2026)
von: Jia, Tianzhi, et al.
Veröffentlicht: (2026)
Fine-grained Image Retrieval via Dual-Vision Adaptation
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
von: Hao, Haihong, et al.
Veröffentlicht: (2025)
von: Hao, Haihong, et al.
Veröffentlicht: (2025)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
von: Tan, Jiawei, et al.
Veröffentlicht: (2024)
von: Tan, Jiawei, et al.
Veröffentlicht: (2024)
SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through Audio-Visual Alignment with Cascaded Cross-Transformer
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
SkyLink: Unifying Street-Satellite Geo-Localization via UAV-Mediated 3D Scene Alignment
von: Zhang, Hongyang, et al.
Veröffentlicht: (2025)
von: Zhang, Hongyang, et al.
Veröffentlicht: (2025)
Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy Labels
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Towards Identity-Aware Cross-Modal Retrieval: a Dataset and a Baseline
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
GeoLink: A 3D-Aware Framework Towards Better Generalization in Cross-View Geo-Localization
von: Zhang, Hongyang, et al.
Veröffentlicht: (2026)
von: Zhang, Hongyang, et al.
Veröffentlicht: (2026)
Multi-scale Activation, Refinement, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird Recognition
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2025)
Joint Explicit and Implicit Cross-Modal Interaction Network for Anterior Chamber Inflammation Diagnosis
von: Shao, Qian, et al.
Veröffentlicht: (2023)
von: Shao, Qian, et al.
Veröffentlicht: (2023)
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
Provably Secure Robust Image Steganography via Cross-Modal Error Correction
von: Qi, Yuang, et al.
Veröffentlicht: (2024)
von: Qi, Yuang, et al.
Veröffentlicht: (2024)
Spatial-Aware Efficient Projector for MLLMs via Multi-Layer Feature Aggregation
von: Qian, Shun, et al.
Veröffentlicht: (2024)
von: Qian, Shun, et al.
Veröffentlicht: (2024)
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
PetalView: Fine-grained Location and Orientation Extraction of Street-view Images via Cross-view Local Search with Supplementary Materials
von: Hu, Wenmiao, et al.
Veröffentlicht: (2024)
von: Hu, Wenmiao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025) -
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
von: Huang, Hailang, et al.
Veröffentlicht: (2024) -
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
von: Chen, Yaru, et al.
Veröffentlicht: (2025) -
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024) -
Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
von: Wang, Haoming, et al.
Veröffentlicht: (2025)