Improved Baselines with Visual Instruction Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Haotian, Li, Chunyuan, Li, Yuheng, Lee, Yong Jae |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Instruct for Visual Instruction Tuning
by: Zhou, Zhihan, et al.
Published: (2025)
by: Zhou, Zhihan, et al.
Published: (2025)
Parrot: Multilingual Visual Instruction Tuning
by: Sun, Hai-Long, et al.
Published: (2024)
by: Sun, Hai-Long, et al.
Published: (2024)
Reconstructive Visual Instruction Tuning
by: Wang, Haochen, et al.
Published: (2024)
by: Wang, Haochen, et al.
Published: (2024)
Leveraging Large Language Models for Scalable Vector Graphics-Driven Image Understanding
by: Cai, Mu, et al.
Published: (2023)
by: Cai, Mu, et al.
Published: (2023)
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
by: Cai, Mu, et al.
Published: (2023)
by: Cai, Mu, et al.
Published: (2023)
Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models
by: Li, Shengzhi, et al.
Published: (2024)
by: Li, Shengzhi, et al.
Published: (2024)
Dr-LLaVA: Visual Instruction Tuning with Symbolic Clinical Grounding
by: Sun, Shenghuan, et al.
Published: (2024)
by: Sun, Shenghuan, et al.
Published: (2024)
LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
by: Fujitake, Masato
Published: (2024)
by: Fujitake, Masato
Published: (2024)
Benchmarking and Analyzing Generative Data for Visual Recognition
by: Li, Bo, et al.
Published: (2023)
by: Li, Bo, et al.
Published: (2023)
MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
by: Lu, Pan, et al.
Published: (2023)
by: Lu, Pan, et al.
Published: (2023)
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
by: Zhang, Jianrui, et al.
Published: (2024)
by: Zhang, Jianrui, et al.
Published: (2024)
Cream of the Crop: Harvesting Rich, Scalable and Transferable Multi-Modal Data for Instruction Fine-Tuning
by: Lyu, Mengyao, et al.
Published: (2025)
by: Lyu, Mengyao, et al.
Published: (2025)
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
by: Zhang, Jianrui, et al.
Published: (2024)
by: Zhang, Jianrui, et al.
Published: (2024)
Matryoshka Multimodal Models
by: Cai, Mu, et al.
Published: (2024)
by: Cai, Mu, et al.
Published: (2024)
Relational Visual Similarity
by: Nguyen, Thao, et al.
Published: (2025)
by: Nguyen, Thao, et al.
Published: (2025)
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
by: Yu, Zhuoran, et al.
Published: (2023)
by: Yu, Zhuoran, et al.
Published: (2023)
Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
by: Qin, Yulei, et al.
Published: (2025)
by: Qin, Yulei, et al.
Published: (2025)
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
by: Li, Baiqi, et al.
Published: (2024)
by: Li, Baiqi, et al.
Published: (2024)
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
by: Pan, Zhiyu, et al.
Published: (2026)
by: Pan, Zhiyu, et al.
Published: (2026)
Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2024)
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2024)
LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation
by: Lee, Suhyeon, et al.
Published: (2023)
by: Lee, Suhyeon, et al.
Published: (2023)
FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering
by: Huang, Chengyue, et al.
Published: (2025)
by: Huang, Chengyue, et al.
Published: (2025)
ReVision: A Dataset and Baseline VLM for Privacy-Preserving Task-Oriented Visual Instruction Rewriting
by: Mishra, Abhijit, et al.
Published: (2025)
by: Mishra, Abhijit, et al.
Published: (2025)
Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning
by: Safaei, Bardia, et al.
Published: (2025)
by: Safaei, Bardia, et al.
Published: (2025)
Real Deep Research for AI, Robotics and Beyond
by: Zou, Xueyan, et al.
Published: (2025)
by: Zou, Xueyan, et al.
Published: (2025)
Vision-Language Models Can Self-Improve Reasoning via Reflection
by: Cheng, Kanzhi, et al.
Published: (2024)
by: Cheng, Kanzhi, et al.
Published: (2024)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
by: Yu, Zhuoran, et al.
Published: (2025)
by: Yu, Zhuoran, et al.
Published: (2025)
Universal Approximation of Visual Autoregressive Transformers
by: Chen, Yifang, et al.
Published: (2025)
by: Chen, Yifang, et al.
Published: (2025)
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
by: Miyai, Atsuyuki, et al.
Published: (2025)
by: Miyai, Atsuyuki, et al.
Published: (2025)
AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance
by: He, Lixuan, et al.
Published: (2025)
by: He, Lixuan, et al.
Published: (2025)
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
by: Yoon, Hyungjun, et al.
Published: (2024)
by: Yoon, Hyungjun, et al.
Published: (2024)
Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Divide & Bind Your Attention for Improved Generative Semantic Nursing
by: Li, Yumeng, et al.
Published: (2023)
by: Li, Yumeng, et al.
Published: (2023)
Exploring Diverse Methods in Visual Question Answering
by: Li, Panfeng, et al.
Published: (2024)
by: Li, Panfeng, et al.
Published: (2024)
ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure
by: Yoon, Hee Suk, et al.
Published: (2023)
by: Yoon, Hee Suk, et al.
Published: (2023)
Instruction-Guided Visual Masking
by: Zheng, Jinliang, et al.
Published: (2024)
by: Zheng, Jinliang, et al.
Published: (2024)
Visual Planning: Let's Think Only with Images
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
Unified Lexical Representation for Interpretable Visual-Language Alignment
by: Li, Yifan, et al.
Published: (2024)
by: Li, Yifan, et al.
Published: (2024)
Similar Items
-
Learning to Instruct for Visual Instruction Tuning
by: Zhou, Zhihan, et al.
Published: (2025) -
Parrot: Multilingual Visual Instruction Tuning
by: Sun, Hai-Long, et al.
Published: (2024) -
Reconstructive Visual Instruction Tuning
by: Wang, Haochen, et al.
Published: (2024) -
Leveraging Large Language Models for Scalable Vector Graphics-Driven Image Understanding
by: Cai, Mu, et al.
Published: (2023) -
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
by: Cai, Mu, et al.
Published: (2023)