Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Juncheng, Pan, Kaihang, Ge, Zhiqi, Gao, Minghe, Ji, Wei, Zhang, Wenqiao, Chua, Tat-Seng, Tang, Siliang, Zhang, Hanwang, Zhuang, Yueting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Auto-Encoding Morph-Tokens for Multimodal LLM
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024)
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024)
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
von: Qiu, Haiyi, et al.
Veröffentlicht: (2024)
von: Qiu, Haiyi, et al.
Veröffentlicht: (2024)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
von: Fan, Zhaoyu, et al.
Veröffentlicht: (2025)
von: Fan, Zhaoyu, et al.
Veröffentlicht: (2025)
Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
von: Gao, Minghe, et al.
Veröffentlicht: (2024)
von: Gao, Minghe, et al.
Veröffentlicht: (2024)
Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
von: Gao, Minghe, et al.
Veröffentlicht: (2025)
von: Gao, Minghe, et al.
Veröffentlicht: (2025)
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
von: Qian, Long, et al.
Veröffentlicht: (2024)
von: Qian, Long, et al.
Veröffentlicht: (2024)
I3: Intent-Introspective Retrieval Conditioned on Instructions
von: Pan, Kaihang, et al.
Veröffentlicht: (2023)
von: Pan, Kaihang, et al.
Veröffentlicht: (2023)
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
von: Qiu, Haiyi, et al.
Veröffentlicht: (2026)
von: Qiu, Haiyi, et al.
Veröffentlicht: (2026)
WorldGPT: Empowering LLM as Multimodal World Model
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024)
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024)
Robust Fine-tuning of Zero-shot Models via Variance Reduction
von: Zhu, Beier, et al.
Veröffentlicht: (2024)
von: Zhu, Beier, et al.
Veröffentlicht: (2024)
Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark
von: Miao, Bingchen, et al.
Veröffentlicht: (2025)
von: Miao, Bingchen, et al.
Veröffentlicht: (2025)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
von: Fei, Hao, et al.
Veröffentlicht: (2023)
von: Fei, Hao, et al.
Veröffentlicht: (2023)
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
von: Gao, Minghe, et al.
Veröffentlicht: (2023)
von: Gao, Minghe, et al.
Veröffentlicht: (2023)
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
von: Yu, Qifan, et al.
Veröffentlicht: (2024)
von: Yu, Qifan, et al.
Veröffentlicht: (2024)
InstructSAM: Segment Any Instance with Any Instructions
von: Yuan, Yuqian, et al.
Veröffentlicht: (2026)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2026)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024)
von: Chow, Wei, et al.
Veröffentlicht: (2024)
OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
von: Bu, Wendong, et al.
Veröffentlicht: (2025)
von: Bu, Wendong, et al.
Veröffentlicht: (2025)
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
von: Cao, Jie, et al.
Veröffentlicht: (2025)
von: Cao, Jie, et al.
Veröffentlicht: (2025)
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
von: Bu, Wendong, et al.
Veröffentlicht: (2025)
von: Bu, Wendong, et al.
Veröffentlicht: (2025)
Towards Semantic Equivalence of Tokenization in Multimodal LLM
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
XNLP: An Interactive Demonstration System for Universal Structured NLP
von: Fei, Hao, et al.
Veröffentlicht: (2023)
von: Fei, Hao, et al.
Veröffentlicht: (2023)
IDEAL: Leveraging Infinite and Dynamic Characterizations of Large Language Models for Query-focused Summarization
von: Cao, Jie, et al.
Veröffentlicht: (2024)
von: Cao, Jie, et al.
Veröffentlicht: (2024)
Bridging Local Details and Global Context in Text-Attributed Graphs
von: Wang, Yaoke, et al.
Veröffentlicht: (2024)
von: Wang, Yaoke, et al.
Veröffentlicht: (2024)
DuetRAG: Collaborative Retrieval-Augmented Generation
von: Jiao, Dian, et al.
Veröffentlicht: (2024)
von: Jiao, Dian, et al.
Veröffentlicht: (2024)
On the Multi-turn Instruction Following for Conversational Web Agents
von: Deng, Yang, et al.
Veröffentlicht: (2024)
von: Deng, Yang, et al.
Veröffentlicht: (2024)
Align$^2$LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation
von: Huang, Hongzhe, et al.
Veröffentlicht: (2024)
von: Huang, Hongzhe, et al.
Veröffentlicht: (2024)
Non-instructional Fine-tuning: Enabling Instruction-Following Capabilities in Pre-trained Language Models without Instruction-Following Data
von: Xie, Juncheng, et al.
Veröffentlicht: (2024)
von: Xie, Juncheng, et al.
Veröffentlicht: (2024)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
von: Chow, Wei, et al.
Veröffentlicht: (2025)
von: Chow, Wei, et al.
Veröffentlicht: (2025)
CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents
von: Wang, Keyu, et al.
Veröffentlicht: (2026)
von: Wang, Keyu, et al.
Veröffentlicht: (2026)
LASER: Tuning-Free LLM-Driven Attention Control for Efficient Text-conditioned Image-to-Animation
von: Zheng, Haoyu, et al.
Veröffentlicht: (2024)
von: Zheng, Haoyu, et al.
Veröffentlicht: (2024)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
von: Wang, Wei, et al.
Veröffentlicht: (2026)
von: Wang, Wei, et al.
Veröffentlicht: (2026)
Data-efficient Fine-tuning for LLM-based Recommendation
von: Lin, Xinyu, et al.
Veröffentlicht: (2024)
von: Lin, Xinyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Auto-Encoding Morph-Tokens for Multimodal LLM
von: Pan, Kaihang, et al.
Veröffentlicht: (2024) -
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
von: Qin, Bosheng, et al.
Veröffentlicht: (2023) -
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024) -
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
von: Qiu, Haiyi, et al.
Veröffentlicht: (2024) -
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
von: Fan, Zhaoyu, et al.
Veröffentlicht: (2025)