Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Jun-Tao, Shi, Yu-Cheng, Xie, Zhen-Hao, Zhou, Da-Wei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning
by: Shi, Yu-Cheng, et al.
Published: (2026)
by: Shi, Yu-Cheng, et al.
Published: (2026)
CRAM: Centroid-Routing and Adaptive MoE for Multimodal Continual Instruction Tuning
by: Tang, Jun-Tao, et al.
Published: (2026)
by: Tang, Jun-Tao, et al.
Published: (2026)
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
by: Hu, Tao, et al.
Published: (2026)
by: Hu, Tao, et al.
Published: (2026)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
by: Chen, Jiaxing, et al.
Published: (2024)
by: Chen, Jiaxing, et al.
Published: (2024)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
by: Guo, Jarvis, et al.
Published: (2024)
by: Guo, Jarvis, et al.
Published: (2024)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
by: Xu, Zhiyang, et al.
Published: (2024)
by: Xu, Zhiyang, et al.
Published: (2024)
SemiHVision: Enhancing Medical Multimodal Models with a Semi-Human Annotated Dataset and Fine-Tuned Instruction Generation
by: Wang, Junda, et al.
Published: (2024)
by: Wang, Junda, et al.
Published: (2024)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
by: Du, Yifan, et al.
Published: (2023)
by: Du, Yifan, et al.
Published: (2023)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024)
by: Qiao, Yuxuan, et al.
Published: (2024)
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
by: Wang, Yanan, et al.
Published: (2025)
by: Wang, Yanan, et al.
Published: (2025)
Expert Pyramid Tuning: Efficient Parameter Fine-Tuning for Expertise-Driven Task Allocation
by: Zhang, Jia-Chen, et al.
Published: (2026)
by: Zhang, Jia-Chen, et al.
Published: (2026)
Parrot: Multilingual Visual Instruction Tuning
by: Sun, Hai-Long, et al.
Published: (2024)
by: Sun, Hai-Long, et al.
Published: (2024)
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
by: Dong, Xuanzhao, et al.
Published: (2026)
by: Dong, Xuanzhao, et al.
Published: (2026)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
by: Zhang, Yuanhan, et al.
Published: (2024)
by: Zhang, Yuanhan, et al.
Published: (2024)
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
by: Tan, Zhiyu, et al.
Published: (2024)
by: Tan, Zhiyu, et al.
Published: (2024)
AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning
by: Xie, Zhen-Hao, et al.
Published: (2026)
by: Xie, Zhen-Hao, et al.
Published: (2026)
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
by: You, Zebin, et al.
Published: (2025)
by: You, Zebin, et al.
Published: (2025)
Less is More: High-value Data Selection for Visual Instruction Tuning
by: Liu, Zikang, et al.
Published: (2024)
by: Liu, Zikang, et al.
Published: (2024)
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
by: Zhou, Li, et al.
Published: (2025)
by: Zhou, Li, et al.
Published: (2025)
Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
by: Zhang, Yanzhe, et al.
Published: (2023)
by: Zhang, Yanzhe, et al.
Published: (2023)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
by: Luo, Chuwei, et al.
Published: (2024)
by: Luo, Chuwei, et al.
Published: (2024)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
by: Wu, Te-Lin, et al.
Published: (2021)
by: Wu, Te-Lin, et al.
Published: (2021)
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
by: Wu, Wei, et al.
Published: (2026)
by: Wu, Wei, et al.
Published: (2026)
Prism: Spectral-Aware Block-Sparse Attention
by: Wang, Xinghao, et al.
Published: (2026)
by: Wang, Xinghao, et al.
Published: (2026)
MANTIS: Interleaved Multi-Image Instruction Tuning
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
by: Guo, Haiyang, et al.
Published: (2025)
by: Guo, Haiyang, et al.
Published: (2025)
Maya: An Instruction Finetuned Multilingual Multimodal Model
by: Alam, Nahid, et al.
Published: (2024)
by: Alam, Nahid, et al.
Published: (2024)
Otter: A Multi-Modal Model with In-Context Instruction Tuning
by: Li, Bo, et al.
Published: (2023)
by: Li, Bo, et al.
Published: (2023)
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
by: Zhang, Wenqiao, et al.
Published: (2024)
by: Zhang, Wenqiao, et al.
Published: (2024)
UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning
by: Tang, Hongxuan, et al.
Published: (2025)
by: Tang, Hongxuan, et al.
Published: (2025)
Unhackable Temporal Rewarding for Scalable Video MLLMs
by: Yu, En, et al.
Published: (2025)
by: Yu, En, et al.
Published: (2025)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
by: Hsieh, Cheng-Yu, et al.
Published: (2025)
by: Hsieh, Cheng-Yu, et al.
Published: (2025)
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
by: Guo, Zirun, et al.
Published: (2024)
by: Guo, Zirun, et al.
Published: (2024)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
by: Nazir, Maham, et al.
Published: (2026)
by: Nazir, Maham, et al.
Published: (2026)
Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
Imp: Highly Capable Large Multimodal Models for Mobile Devices
by: Shao, Zhenwei, et al.
Published: (2024)
by: Shao, Zhenwei, et al.
Published: (2024)
Similar Items
-
ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning
by: Shi, Yu-Cheng, et al.
Published: (2026) -
CRAM: Centroid-Routing and Adaptive MoE for Multimodal Continual Instruction Tuning
by: Tang, Jun-Tao, et al.
Published: (2026) -
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
by: Hu, Tao, et al.
Published: (2026) -
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
by: Chen, Jiaxing, et al.
Published: (2024) -
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
by: Guo, Jarvis, et al.
Published: (2024)