Saved in:
| Main Authors: | Pang, Yijiang, Hoang, Bao, Zhou, Jiayu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2403.07888 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
by: Shi, Danli, et al.
Published: (2024)
by: Shi, Danli, et al.
Published: (2024)
Cross-modal linkage risk in clinical vision-language models
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval
by: Lei, Youbo, et al.
Published: (2023)
by: Lei, Youbo, et al.
Published: (2023)
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
by: Li, Bing, et al.
Published: (2022)
by: Li, Bing, et al.
Published: (2022)
NCSAM Noise-Compensated Sharpness-Aware Minimization for Noisy Label Learning
by: Xu, Jiayu, et al.
Published: (2026)
by: Xu, Jiayu, et al.
Published: (2026)
Masked Contrastive Reconstruction for Cross-modal Medical Image-Report Retrieval
by: Wei, Zeqiang, et al.
Published: (2023)
by: Wei, Zeqiang, et al.
Published: (2023)
Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification
by: Liu, Tengfei, et al.
Published: (2024)
by: Liu, Tengfei, et al.
Published: (2024)
Mitigating annotation shift in cancer classification using single image generative models
by: Arcas, Marta Buetas, et al.
Published: (2024)
by: Arcas, Marta Buetas, et al.
Published: (2024)
Fusion-Mamba for Cross-modality Object Detection
by: Dong, Wenhao, et al.
Published: (2024)
by: Dong, Wenhao, et al.
Published: (2024)
Variational Adapter for Cross-modal Similarity Representation
by: Wei, WenZhang, et al.
Published: (2026)
by: Wei, WenZhang, et al.
Published: (2026)
VLA-Mark: A cross modal watermark for large vision-language alignment model
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
Cross-modality Guidance-aided Multi-modal Learning with Dual Attention for MRI Brain Tumor Grading
by: Xu, Dunyuan, et al.
Published: (2024)
by: Xu, Dunyuan, et al.
Published: (2024)
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
by: Chen, Haonan, et al.
Published: (2026)
by: Chen, Haonan, et al.
Published: (2026)
XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation
by: Wang, Ziyi, et al.
Published: (2024)
by: Wang, Ziyi, et al.
Published: (2024)
Where are we with calibration under dataset shift in image classification?
by: Roschewitz, Mélanie, et al.
Published: (2025)
by: Roschewitz, Mélanie, et al.
Published: (2025)
Diffexplainer: Towards Cross-modal Global Explanations with Diffusion Models
by: Pennisi, Matteo, et al.
Published: (2024)
by: Pennisi, Matteo, et al.
Published: (2024)
xModel-KD: Cross-modal Knowledge Distillation for 3D Scene Perception using LiDAR
by: Pathmanathan, Thenukan, et al.
Published: (2026)
by: Pathmanathan, Thenukan, et al.
Published: (2026)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
AgriKD: Cross-Architecture Knowledge Distillation for Efficient Leaf Disease Classification
by: Le, Minh-Dung, et al.
Published: (2026)
by: Le, Minh-Dung, et al.
Published: (2026)
Bridging Modality Gap for Visual Grounding with Effecitve Cross-modal Distillation
by: Wang, Jiaxi, et al.
Published: (2023)
by: Wang, Jiaxi, et al.
Published: (2023)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
Cross-modal Causal Intervention for Alzheimer's Disease Prediction
by: Jin, Yutao, et al.
Published: (2025)
by: Jin, Yutao, et al.
Published: (2025)
Automatic dataset shift identification to support safe deployment of medical imaging AI
by: Roschewitz, Mélanie, et al.
Published: (2024)
by: Roschewitz, Mélanie, et al.
Published: (2024)
SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models
by: Fang, Hengyu, et al.
Published: (2025)
by: Fang, Hengyu, et al.
Published: (2025)
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
Memory-based Cross-modal Semantic Alignment Network for Radiology Report Generation
by: Tao, Yitian, et al.
Published: (2024)
by: Tao, Yitian, et al.
Published: (2024)
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
AnomalyControl: Learning Cross-modal Semantic Features for Controllable Anomaly Synthesis
by: He, Shidan, et al.
Published: (2024)
by: He, Shidan, et al.
Published: (2024)
Benefit from Reference: Retrieval-Augmented Cross-modal Point Cloud Completion
by: Hou, Hongye, et al.
Published: (2025)
by: Hou, Hongye, et al.
Published: (2025)
Cross-domain Few-shot Object Detection with Multi-modal Textual Enrichment
by: Shangguan, Zeyu, et al.
Published: (2025)
by: Shangguan, Zeyu, et al.
Published: (2025)
Multi-modal Relation Distillation for Unified 3D Representation Learning
by: Wang, Huiqun, et al.
Published: (2024)
by: Wang, Huiqun, et al.
Published: (2024)
C3-Diff: Super-resolving Spatial Transcriptomics via Cross-modal Cross-content Contrastive Diffusion Modelling
by: Wang, Xiaofei, et al.
Published: (2025)
by: Wang, Xiaofei, et al.
Published: (2025)
Awesome Multi-modal Object Tracking
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
by: Zhang, Yudong, et al.
Published: (2024)
by: Zhang, Yudong, et al.
Published: (2024)
Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
by: He, Xiao, et al.
Published: (2025)
by: He, Xiao, et al.
Published: (2025)
TACFN: Transformer-based Adaptive Cross-modal Fusion Network for Multimodal Emotion Recognition
by: Liu, Feng, et al.
Published: (2025)
by: Liu, Feng, et al.
Published: (2025)
Cross-modal Information Flow in Multimodal Large Language Models
by: Zhang, Zhi, et al.
Published: (2024)
by: Zhang, Zhi, et al.
Published: (2024)
CleverDistiller: Simple and Spatially Consistent Cross-modal Distillation
by: Govindarajan, Hariprasath, et al.
Published: (2025)
by: Govindarajan, Hariprasath, et al.
Published: (2025)
ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking
by: Ge, Jiawei, et al.
Published: (2026)
by: Ge, Jiawei, et al.
Published: (2026)
Structurally Consistent MRI Colorization using Cross-modal Fusion Learning
by: Mathur, Mayuri, et al.
Published: (2024)
by: Mathur, Mayuri, et al.
Published: (2024)
Similar Items
-
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
by: Shi, Danli, et al.
Published: (2024) -
Cross-modal linkage risk in clinical vision-language models
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026) -
MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval
by: Lei, Youbo, et al.
Published: (2023) -
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
by: Li, Bing, et al.
Published: (2022) -
NCSAM Noise-Compensated Sharpness-Aware Minimization for Noisy Label Learning
by: Xu, Jiayu, et al.
Published: (2026)