MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Haiyang, Zhu, Fei, Zhao, Hongbo, Zeng, Fanhu, Liu, Wenzhuo, Ma, Shijie, Wang, Da-Han, Zhang, Xu-Yao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CL-VISTA: Benchmarking Continual Learning in Video Large Language Models
by: Guo, Haiyang, et al.
Published: (2026)
by: Guo, Haiyang, et al.
Published: (2026)
ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
by: Zeng, Fanhu, et al.
Published: (2024)
by: Zeng, Fanhu, et al.
Published: (2024)
Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
by: Guo, Haiyang, et al.
Published: (2025)
by: Guo, Haiyang, et al.
Published: (2025)
LLaVA-c: Continual Improved Visual Instruction Tuning
by: Liu, Wenzhuo, et al.
Published: (2025)
by: Liu, Wenzhuo, et al.
Published: (2025)
RobustMerge: Parameter-Efficient Model Merging for MLLMs with Direction Robustness
by: Zeng, Fanhu, et al.
Published: (2025)
by: Zeng, Fanhu, et al.
Published: (2025)
Happy: A Debiased Learning Framework for Continual Generalized Category Discovery
by: Ma, Shijie, et al.
Published: (2024)
by: Ma, Shijie, et al.
Published: (2024)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
Federated Continual Instruction Tuning
by: Guo, Haiyang, et al.
Published: (2025)
by: Guo, Haiyang, et al.
Published: (2025)
PILoRA: Prototype Guided Incremental LoRA for Federated Class-Incremental Learning
by: Guo, Haiyang, et al.
Published: (2024)
by: Guo, Haiyang, et al.
Published: (2024)
Towards Efficient and General-Purpose Few-Shot Misclassification Detection for Vision-Language Models
by: Zeng, Fanhu, et al.
Published: (2025)
by: Zeng, Fanhu, et al.
Published: (2025)
Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
Local-Prompt: Extensible Local Prompts for Few-Shot Out-of-Distribution Detection
by: Zeng, Fanhu, et al.
Published: (2024)
by: Zeng, Fanhu, et al.
Published: (2024)
Branch-Tuning: Balancing Stability and Plasticity for Continual Self-Supervised Learning
by: Liu, Wenzhuo, et al.
Published: (2024)
by: Liu, Wenzhuo, et al.
Published: (2024)
MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any Resolution
by: Liu, Wenzhuo, et al.
Published: (2024)
by: Liu, Wenzhuo, et al.
Published: (2024)
MLLM-CL: Continual Learning for Multimodal Large Language Models
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
by: Hu, Tao, et al.
Published: (2026)
by: Hu, Tao, et al.
Published: (2026)
Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
by: Ge, Chendi, et al.
Published: (2025)
by: Ge, Chendi, et al.
Published: (2025)
HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model
by: Guo, Haiyang, et al.
Published: (2025)
by: Guo, Haiyang, et al.
Published: (2025)
MMDG-Bench: A Benchmark for Multimodal Domain Generalization
by: Zhan, Qianshan, et al.
Published: (2026)
by: Zhan, Qianshan, et al.
Published: (2026)
Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning
by: Tang, Jun-Tao, et al.
Published: (2026)
by: Tang, Jun-Tao, et al.
Published: (2026)
ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning
by: Shi, Yu-Cheng, et al.
Published: (2026)
by: Shi, Yu-Cheng, et al.
Published: (2026)
Multimodal Instruction Tuning with Hybrid State Space Models
by: Zhou, Jianing, et al.
Published: (2024)
by: Zhou, Jianing, et al.
Published: (2024)
Multimodal Instruction Tuning with Conditional Mixture of LoRA
by: Shen, Ying, et al.
Published: (2024)
by: Shen, Ying, et al.
Published: (2024)
Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM Tasks
by: Jin, Jing, et al.
Published: (2026)
by: Jin, Jing, et al.
Published: (2026)
Towards Trustworthy Dataset Distillation
by: Ma, Shijie, et al.
Published: (2023)
by: Ma, Shijie, et al.
Published: (2023)
3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset
by: Zhang, Junjie, et al.
Published: (2024)
by: Zhang, Junjie, et al.
Published: (2024)
Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration
by: Zeng, Fanhu, et al.
Published: (2025)
by: Zeng, Fanhu, et al.
Published: (2025)
PPT: Token Pruning and Pooling for Efficient Vision Transformers
by: Wu, Xinjian, et al.
Published: (2023)
by: Wu, Xinjian, et al.
Published: (2023)
Multi-scale Unified Network for Image Classification
by: Liu, Wenzhuo, et al.
Published: (2024)
by: Liu, Wenzhuo, et al.
Published: (2024)
Towards Non-Exemplar Semi-Supervised Class-Incremental Learning
by: Liu, Wenzhuo, et al.
Published: (2024)
by: Liu, Wenzhuo, et al.
Published: (2024)
MERGETUNE: Continued Fine-Tuning of Vision-Language Models
by: Wang, Wenqing, et al.
Published: (2026)
by: Wang, Wenqing, et al.
Published: (2026)
Active Generalized Category Discovery
by: Ma, Shijie, et al.
Published: (2024)
by: Ma, Shijie, et al.
Published: (2024)
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuning
by: Yao, Zhenquan, et al.
Published: (2026)
by: Yao, Zhenquan, et al.
Published: (2026)
CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
by: Chen, Cheng, et al.
Published: (2024)
by: Chen, Cheng, et al.
Published: (2024)
DiffVC-RT: Towards Practical Real-Time Diffusion-based Perceptual Neural Video Compression
by: Ma, Wenzhuo, et al.
Published: (2026)
by: Ma, Wenzhuo, et al.
Published: (2026)
Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning
by: Zeng, Xingchen, et al.
Published: (2024)
by: Zeng, Xingchen, et al.
Published: (2024)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
by: Guo, Jarvis, et al.
Published: (2024)
by: Guo, Jarvis, et al.
Published: (2024)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection
by: Zhao, Haochen, et al.
Published: (2025)
by: Zhao, Haochen, et al.
Published: (2025)
Similar Items
-
CL-VISTA: Benchmarking Continual Learning in Video Large Language Models
by: Guo, Haiyang, et al.
Published: (2026) -
ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
by: Zeng, Fanhu, et al.
Published: (2024) -
Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
by: Guo, Haiyang, et al.
Published: (2025) -
LLaVA-c: Continual Improved Visual Instruction Tuning
by: Liu, Wenzhuo, et al.
Published: (2025) -
RobustMerge: Parameter-Efficient Model Merging for MLLMs with Direction Robustness
by: Zeng, Fanhu, et al.
Published: (2025)