Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Lincan, Li, Shuang, Ma, Wenxuan, Kang, Jingxuan, Xie, Binhui, Sun, Zixun, Zhu, Chengwei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Modality Knowledge Alignment for Cross-Modality Transfer
by: Ma, Wenxuan, et al.
Published: (2024)
by: Ma, Wenxuan, et al.
Published: (2024)
From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
by: Cai, Lincan, et al.
Published: (2025)
by: Cai, Lincan, et al.
Published: (2025)
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
by: Hu, Tao, et al.
Published: (2026)
by: Hu, Tao, et al.
Published: (2026)
Unsupervised Modality-Transferable Video Highlight Detection with Representation Activation Sequence Learning
by: Li, Tingtian, et al.
Published: (2024)
by: Li, Tingtian, et al.
Published: (2024)
Bayesian Cross-Modal Alignment Learning for Few-Shot Out-of-Distribution Generalization
by: Zhu, Lin, et al.
Published: (2025)
by: Zhu, Lin, et al.
Published: (2025)
ModalTune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-task Learning in Digital Pathology
by: Ramanathan, Vishwesh, et al.
Published: (2025)
by: Ramanathan, Vishwesh, et al.
Published: (2025)
Cross-Modal Few-Shot Learning: a Generative Transfer Learning Framework
by: Yang, Zhengwei, et al.
Published: (2024)
by: Yang, Zhengwei, et al.
Published: (2024)
On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
by: Cai, Yichao, et al.
Published: (2025)
by: Cai, Yichao, et al.
Published: (2025)
Cross-Modal Adapter: Parameter-Efficient Transfer Learning Approach for Vision-Language Models
by: Yang, Juncheng, et al.
Published: (2024)
by: Yang, Juncheng, et al.
Published: (2024)
pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models
by: Ghiasvand, Sajjad, et al.
Published: (2025)
by: Ghiasvand, Sajjad, et al.
Published: (2025)
A Generalization Theory of Cross-Modality Distillation with Contrastive Learning
by: Lin, Hangyu, et al.
Published: (2024)
by: Lin, Hangyu, et al.
Published: (2024)
Mind the Gap: Learning Modality-Agnostic Representations with a Cross-Modality UNet
by: Niu, Xin, et al.
Published: (2026)
by: Niu, Xin, et al.
Published: (2026)
Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
by: Naharas, Nilay, et al.
Published: (2025)
by: Naharas, Nilay, et al.
Published: (2025)
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
by: Jiang, Lei, et al.
Published: (2025)
by: Jiang, Lei, et al.
Published: (2025)
Conformal Cross-Modal Active Learning
by: Nguyen, Huy Hoang, et al.
Published: (2026)
by: Nguyen, Huy Hoang, et al.
Published: (2026)
Cross-Modal Coordination Across a Diverse Set of Input Modalities
by: Sánchez, Jorge, et al.
Published: (2024)
by: Sánchez, Jorge, et al.
Published: (2024)
Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization
by: Wang, Xiaohan, et al.
Published: (2025)
by: Wang, Xiaohan, et al.
Published: (2025)
CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation
by: Jeon, Inseok, et al.
Published: (2026)
by: Jeon, Inseok, et al.
Published: (2026)
PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities
by: Chen, Jiajun, et al.
Published: (2025)
by: Chen, Jiajun, et al.
Published: (2025)
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
by: Li, Xu, et al.
Published: (2025)
by: Li, Xu, et al.
Published: (2025)
Modality Unified Attack for Omni-Modality Person Re-Identification
by: Bian, Yuan, et al.
Published: (2025)
by: Bian, Yuan, et al.
Published: (2025)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
by: Fragomeni, Adriano, et al.
Published: (2025)
by: Fragomeni, Adriano, et al.
Published: (2025)
Cross-Modal Instructions for Robot Motion Generation
by: Barron, William, et al.
Published: (2025)
by: Barron, William, et al.
Published: (2025)
CM2-Net: Continual Cross-Modal Mapping Network for Driver Action Recognition
by: Wang, Ruoyu, et al.
Published: (2024)
by: Wang, Ruoyu, et al.
Published: (2024)
Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality
by: Wang, Hu, et al.
Published: (2023)
by: Wang, Hu, et al.
Published: (2023)
MUST: Modality-Specific Representation-Aware Transformer for Diffusion-Enhanced Survival Prediction with Missing Modality
by: Kim, Kyungwon, et al.
Published: (2026)
by: Kim, Kyungwon, et al.
Published: (2026)
MolFM-Lite: Multi-Modal Molecular Property Prediction with Conformer Ensemble Attention and Cross-Modal Fusion
by: Shah, Syed Omer, et al.
Published: (2026)
by: Shah, Syed Omer, et al.
Published: (2026)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings
by: Dhimoïla, Grégoire, et al.
Published: (2026)
by: Dhimoïla, Grégoire, et al.
Published: (2026)
Quantifying Cross-Modality Memorization in Vision-Language Models
by: Wen, Yuxin, et al.
Published: (2025)
by: Wen, Yuxin, et al.
Published: (2025)
Continual Cross-Modal Generalization
by: Xia, Yan, et al.
Published: (2025)
by: Xia, Yan, et al.
Published: (2025)
Cross-Modal Domain Adaptation in Brain Disease Diagnosis: Maximum Mean Discrepancy-based Convolutional Neural Networks
by: Zhu, Xuran
Published: (2024)
by: Zhu, Xuran
Published: (2024)
Balanced Multi-modal Federated Learning via Cross-Modal Infiltration
by: Fan, Yunfeng, et al.
Published: (2023)
by: Fan, Yunfeng, et al.
Published: (2023)
Cross-Modal Fine-Tuning of 3D Convolutional Foundation Models for ADHD Classification with Low-Rank Adaptation
by: Kao, Jyun-Ping, et al.
Published: (2025)
by: Kao, Jyun-Ping, et al.
Published: (2025)
Text-to-Image Cross-Modal Generation: A Systematic Review
by: Żelaszczyk, Maciej, et al.
Published: (2024)
by: Żelaszczyk, Maciej, et al.
Published: (2024)
Cross-Modal Clinical Knowledge Integration for Mammography Report Generation
by: Zhu, Jiayi, et al.
Published: (2026)
by: Zhu, Jiayi, et al.
Published: (2026)
Lightweight Facial Landmark Detection in Thermal Images via Multi-Level Cross-Modal Knowledge Transfer
by: Tong, Qiyi, et al.
Published: (2025)
by: Tong, Qiyi, et al.
Published: (2025)
CRAFT: Aligning Diffusion Models with Fine-Tuning Is Easier Than You Think
by: Sun, Zening, et al.
Published: (2026)
by: Sun, Zening, et al.
Published: (2026)
Cross-Modal Mapping: Mitigating the Modality Gap for Few-Shot Image Classification
by: Yang, Xi, et al.
Published: (2024)
by: Yang, Xi, et al.
Published: (2024)
Enhancing CLIP Robustness via Cross-Modality Alignment
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
Similar Items
-
Learning Modality Knowledge Alignment for Cross-Modality Transfer
by: Ma, Wenxuan, et al.
Published: (2024) -
From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
by: Cai, Lincan, et al.
Published: (2025) -
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
by: Hu, Tao, et al.
Published: (2026) -
Unsupervised Modality-Transferable Video Highlight Detection with Representation Activation Sequence Learning
by: Li, Tingtian, et al.
Published: (2024) -
Bayesian Cross-Modal Alignment Learning for Few-Shot Out-of-Distribution Generalization
by: Zhu, Lin, et al.
Published: (2025)