MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shan, Xiaojun, Cao, Qi, Han, Xing, Yu, Haofei, Liang, Paul Pu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026)
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026)
MAny: Merge Anything for Multimodal Continual Instruction Tuning
von: Gao, Zijian, et al.
Veröffentlicht: (2026)
von: Gao, Zijian, et al.
Veröffentlicht: (2026)
Understanding the Emergence of Multimodal Representation Alignment
von: Tjandrasuwita, Megan, et al.
Veröffentlicht: (2025)
von: Tjandrasuwita, Megan, et al.
Veröffentlicht: (2025)
HEMM: Holistic Evaluation of Multimodal Foundation Models
von: Liang, Paul Pu, et al.
Veröffentlicht: (2024)
von: Liang, Paul Pu, et al.
Veröffentlicht: (2024)
Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
von: Ryan, Yuriel, et al.
Veröffentlicht: (2026)
von: Ryan, Yuriel, et al.
Veröffentlicht: (2026)
SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning
von: Chen, Jinpeng, et al.
Veröffentlicht: (2025)
von: Chen, Jinpeng, et al.
Veröffentlicht: (2025)
CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
von: Wang, Yongxin, et al.
Veröffentlicht: (2025)
von: Wang, Yongxin, et al.
Veröffentlicht: (2025)
M$^2$PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning
von: Wang, Taowen, et al.
Veröffentlicht: (2024)
von: Wang, Taowen, et al.
Veröffentlicht: (2024)
MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer's Screening
von: Ahire, Vrushank, et al.
Veröffentlicht: (2026)
von: Ahire, Vrushank, et al.
Veröffentlicht: (2026)
Adapt-$\infty$: Scalable Continual Multimodal Instruction Tuning via Dynamic Data Selection
von: Maharana, Adyasha, et al.
Veröffentlicht: (2024)
von: Maharana, Adyasha, et al.
Veröffentlicht: (2024)
Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
von: Dima, George-Andrei, et al.
Veröffentlicht: (2025)
von: Dima, George-Andrei, et al.
Veröffentlicht: (2025)
Pilot: Building the Federated Multimodal Instruction Tuning Framework
von: Xiong, Baochen, et al.
Veröffentlicht: (2025)
von: Xiong, Baochen, et al.
Veröffentlicht: (2025)
Multimodal Fine-grained Reasoning for Post Quality Evaluation
von: Guo, Xiaoxu, et al.
Veröffentlicht: (2025)
von: Guo, Xiaoxu, et al.
Veröffentlicht: (2025)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
von: Wang, Xingyao, et al.
Veröffentlicht: (2023)
von: Wang, Xingyao, et al.
Veröffentlicht: (2023)
MultiMed: Massively Multimodal and Multitask Medical Understanding
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
Revealing Multimodal Causality with Large Language Models
von: Li, Jin, et al.
Veröffentlicht: (2025)
von: Li, Jin, et al.
Veröffentlicht: (2025)
Multimodal Web Navigation with Instruction-Finetuned Foundation Models
von: Furuta, Hiroki, et al.
Veröffentlicht: (2023)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2023)
MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data
von: Zhu, Zhenghao, et al.
Veröffentlicht: (2025)
von: Zhu, Zhenghao, et al.
Veröffentlicht: (2025)
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
von: Hu, Tao, et al.
Veröffentlicht: (2026)
von: Hu, Tao, et al.
Veröffentlicht: (2026)
Federated Continual Instruction Tuning
von: Guo, Haiyang, et al.
Veröffentlicht: (2025)
von: Guo, Haiyang, et al.
Veröffentlicht: (2025)
Discrete Diffusion in Large Language and Multimodal Models: A Survey
von: Yu, Runpeng, et al.
Veröffentlicht: (2025)
von: Yu, Runpeng, et al.
Veröffentlicht: (2025)
Generative Representational Instruction Tuning
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
Efficient Quantification of Multimodal Interaction at Sample Level
von: Yang, Zequn, et al.
Veröffentlicht: (2025)
von: Yang, Zequn, et al.
Veröffentlicht: (2025)
Feature-level Interaction Explanations in Multimodal Transformers
von: Kim, Yeji, et al.
Veröffentlicht: (2026)
von: Kim, Yeji, et al.
Veröffentlicht: (2026)
Multimodal Negative Learning
von: Gong, Baoquan, et al.
Veröffentlicht: (2025)
von: Gong, Baoquan, et al.
Veröffentlicht: (2025)
QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
von: Dai, Wei, et al.
Veröffentlicht: (2025)
von: Dai, Wei, et al.
Veröffentlicht: (2025)
Neuro-Inspired Hierarchical Multimodal Learning
von: Xiao, Xiongye, et al.
Veröffentlicht: (2023)
von: Xiao, Xiongye, et al.
Veröffentlicht: (2023)
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
von: Cao, Qi, et al.
Veröffentlicht: (2025)
von: Cao, Qi, et al.
Veröffentlicht: (2025)
FAIRWELL: Fair Multimodal Self-Supervised Learning for Wellbeing Prediction
von: Cheong, Jiaee, et al.
Veröffentlicht: (2025)
von: Cheong, Jiaee, et al.
Veröffentlicht: (2025)
Read to Play (R2-Play): Decision Transformer with Multimodal Game Instruction
von: Jin, Yonggang, et al.
Veröffentlicht: (2024)
von: Jin, Yonggang, et al.
Veröffentlicht: (2024)
Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models
von: Jia, Yifan, et al.
Veröffentlicht: (2025)
von: Jia, Yifan, et al.
Veröffentlicht: (2025)
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
MMCTOP: A Multimodal Textualization and Mixture-of-Experts Framework for Clinical Trial Outcome Prediction
von: Aparício, Carolina, et al.
Veröffentlicht: (2025)
von: Aparício, Carolina, et al.
Veröffentlicht: (2025)
Multimodal Lego: Model Merging and Fine-Tuning Across Topologies and Modalities in Biomedicine
von: Hemker, Konstantin, et al.
Veröffentlicht: (2024)
von: Hemker, Konstantin, et al.
Veröffentlicht: (2024)
What Makes a Good Diffusion Planner for Decision Making?
von: Lu, Haofei, et al.
Veröffentlicht: (2025)
von: Lu, Haofei, et al.
Veröffentlicht: (2025)
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025)
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025)
Learning Multimodal Energy-Based Model with Multimodal Variational Auto-Encoder via MCMC Revision
von: Cui, Jiali, et al.
Veröffentlicht: (2026)
von: Cui, Jiali, et al.
Veröffentlicht: (2026)
Diffusion Instruction Tuning
von: Jin, Chen, et al.
Veröffentlicht: (2025)
von: Jin, Chen, et al.
Veröffentlicht: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026) -
MAny: Merge Anything for Multimodal Continual Instruction Tuning
von: Gao, Zijian, et al.
Veröffentlicht: (2026) -
Understanding the Emergence of Multimodal Representation Alignment
von: Tjandrasuwita, Megan, et al.
Veröffentlicht: (2025) -
HEMM: Holistic Evaluation of Multimodal Foundation Models
von: Liang, Paul Pu, et al.
Veröffentlicht: (2024) -
Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
von: Ryan, Yuriel, et al.
Veröffentlicht: (2026)