Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Tao, Zhou, Da-Wei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
by: Li, Xu, et al.
Published: (2025)
by: Li, Xu, et al.
Published: (2025)
Cross-Modal Instructions for Robot Motion Generation
by: Barron, William, et al.
Published: (2025)
by: Barron, William, et al.
Published: (2025)
Pilot: Building the Federated Multimodal Instruction Tuning Framework
by: Xiong, Baochen, et al.
Published: (2025)
by: Xiong, Baochen, et al.
Published: (2025)
Adaptive Prompt Tuning: Vision Guided Prompt Tuning with Cross-Attention for Fine-Grained Few-Shot Learning
by: Brouwer, Eric, et al.
Published: (2024)
by: Brouwer, Eric, et al.
Published: (2024)
Parrot: Multilingual Visual Instruction Tuning
by: Sun, Hai-Long, et al.
Published: (2024)
by: Sun, Hai-Long, et al.
Published: (2024)
Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning
by: Tang, Jun-Tao, et al.
Published: (2026)
by: Tang, Jun-Tao, et al.
Published: (2026)
ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning
by: Shi, Yu-Cheng, et al.
Published: (2026)
by: Shi, Yu-Cheng, et al.
Published: (2026)
Robust Multimodal Learning via Cross-Modal Proxy Tokens
by: Reza, Md Kaykobad, et al.
Published: (2025)
by: Reza, Md Kaykobad, et al.
Published: (2025)
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
by: Ai, Jiaxin, et al.
Published: (2025)
by: Ai, Jiaxin, et al.
Published: (2025)
Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning
by: Safaei, Bardia, et al.
Published: (2025)
by: Safaei, Bardia, et al.
Published: (2025)
DiPrompT: Disentangled Prompt Tuning for Multiple Latent Domain Generalization in Federated Learning
by: Bai, Sikai, et al.
Published: (2024)
by: Bai, Sikai, et al.
Published: (2024)
Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks
by: Matsuishi, Koki, et al.
Published: (2025)
by: Matsuishi, Koki, et al.
Published: (2025)
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
by: Li, Bingyu, et al.
Published: (2024)
by: Li, Bingyu, et al.
Published: (2024)
Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs
by: Hu, Zixuan, et al.
Published: (2024)
by: Hu, Zixuan, et al.
Published: (2024)
Deep Multimodal Learning with Missing Modality: A Survey
by: Wu, Renjie, et al.
Published: (2024)
by: Wu, Renjie, et al.
Published: (2024)
MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
by: Guo, Haiyang, et al.
Published: (2025)
by: Guo, Haiyang, et al.
Published: (2025)
Cream of the Crop: Harvesting Rich, Scalable and Transferable Multi-Modal Data for Instruction Fine-Tuning
by: Lyu, Mengyao, et al.
Published: (2025)
by: Lyu, Mengyao, et al.
Published: (2025)
MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
by: Liu, Ziyu, et al.
Published: (2024)
by: Liu, Ziyu, et al.
Published: (2024)
Distilled Prompt Learning for Incomplete Multimodal Survival Prediction
by: Xu, Yingxue, et al.
Published: (2025)
by: Xu, Yingxue, et al.
Published: (2025)
Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models
by: Song, Fei, et al.
Published: (2025)
by: Song, Fei, et al.
Published: (2025)
Learning to Instruct for Visual Instruction Tuning
by: Zhou, Zhihan, et al.
Published: (2025)
by: Zhou, Zhihan, et al.
Published: (2025)
CroMe: Multimodal Fake News Detection using Cross-Modal Tri-Transformer and Metric Learning
by: Choi, Eunjee, et al.
Published: (2025)
by: Choi, Eunjee, et al.
Published: (2025)
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
FlyPrompt: Brain-Inspired Random-Expanded Routing with Temporal-Ensemble Experts for General Continual Learning
by: Yan, Hongwei, et al.
Published: (2026)
by: Yan, Hongwei, et al.
Published: (2026)
Candidate Pseudolabel Learning: Enhancing Vision-Language Models by Prompt Tuning with Unlabeled Data
by: Zhang, Jiahan, et al.
Published: (2024)
by: Zhang, Jiahan, et al.
Published: (2024)
ADAPT to Robustify Prompt Tuning Vision Transformers
by: Eskandar, Masih, et al.
Published: (2024)
by: Eskandar, Masih, et al.
Published: (2024)
Unified Continuous Generative Models
by: Sun, Peng, et al.
Published: (2025)
by: Sun, Peng, et al.
Published: (2025)
RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer
by: Ni, Haotian, et al.
Published: (2025)
by: Ni, Haotian, et al.
Published: (2025)
RainbowPrompt: Diversity-Enhanced Prompt-Evolving for Continual Learning
by: Hong, Kiseong, et al.
Published: (2025)
by: Hong, Kiseong, et al.
Published: (2025)
Continual Adapter Tuning with Semantic Shift Compensation for Class-Incremental Learning
by: Zhou, Qinhao, et al.
Published: (2024)
by: Zhou, Qinhao, et al.
Published: (2024)
Reconstructive Visual Instruction Tuning
by: Wang, Haochen, et al.
Published: (2024)
by: Wang, Haochen, et al.
Published: (2024)
Visual Modality Prompt for Adapting Vision-Language Object Detectors
by: Medeiros, Heitor R., et al.
Published: (2024)
by: Medeiros, Heitor R., et al.
Published: (2024)
Cross-Source Supervision for Bone Infection Segmentation in Dual-Modality PET-CT
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
by: Yaras, Can, et al.
Published: (2024)
by: Yaras, Can, et al.
Published: (2024)
Improved Baselines with Visual Instruction Tuning
by: Liu, Haotian, et al.
Published: (2023)
by: Liu, Haotian, et al.
Published: (2023)
PackDiT: Joint Human Motion and Text Generation via Mutual Prompting
by: Jiang, Zhongyu, et al.
Published: (2025)
by: Jiang, Zhongyu, et al.
Published: (2025)
CLASH: A Benchmark for Cross-Modal Contradiction Detection
by: Popordanoska, Teodora, et al.
Published: (2025)
by: Popordanoska, Teodora, et al.
Published: (2025)
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
GIST: Targeted Data Selection for Instruction Tuning via Coupled Optimization Geometry
by: Min, Guanghui, et al.
Published: (2026)
by: Min, Guanghui, et al.
Published: (2026)
MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
Similar Items
-
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
by: Li, Xu, et al.
Published: (2025) -
Cross-Modal Instructions for Robot Motion Generation
by: Barron, William, et al.
Published: (2025) -
Pilot: Building the Federated Multimodal Instruction Tuning Framework
by: Xiong, Baochen, et al.
Published: (2025) -
Adaptive Prompt Tuning: Vision Guided Prompt Tuning with Cross-Attention for Fine-Grained Few-Shot Learning
by: Brouwer, Eric, et al.
Published: (2024) -
Parrot: Multilingual Visual Instruction Tuning
by: Sun, Hai-Long, et al.
Published: (2024)