What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Du, Yifan, Guo, Hangyu, Zhou, Kun, Zhao, Wayne Xin, Wang, Jinpeng, Wang, Chuyuan, Cai, Mingchen, Song, Ruihua, Wen, Ji-Rong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Less is More: High-value Data Selection for Visual Instruction Tuning
von: Liu, Zikang, et al.
Veröffentlicht: (2024)
von: Liu, Zikang, et al.
Veröffentlicht: (2024)
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
Reconstructive Visual Instruction Tuning
von: Wang, Haochen, et al.
Veröffentlicht: (2024)
von: Wang, Haochen, et al.
Veröffentlicht: (2024)
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
von: Li, Yifan, et al.
Veröffentlicht: (2024)
von: Li, Yifan, et al.
Veröffentlicht: (2024)
EAGLE: Elevating Geometric Reasoning through LLM-empowered Visual Instruction Tuning
von: Li, Zhihao, et al.
Veröffentlicht: (2024)
von: Li, Zhihao, et al.
Veröffentlicht: (2024)
Learning to Instruct for Visual Instruction Tuning
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
Personalized Visual Instruction Tuning
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
Generative Visual Instruction Tuning
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2024)
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2024)
Comparison Visual Instruction Tuning
von: Lin, Wei, et al.
Veröffentlicht: (2024)
von: Lin, Wei, et al.
Veröffentlicht: (2024)
Connecting Dreams with Visual Brainstorming Instruction
von: Sun, Yasheng, et al.
Veröffentlicht: (2024)
von: Sun, Yasheng, et al.
Veröffentlicht: (2024)
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
von: You, Zebin, et al.
Veröffentlicht: (2025)
von: You, Zebin, et al.
Veröffentlicht: (2025)
Improved Baselines with Visual Instruction Tuning
von: Liu, Haotian, et al.
Veröffentlicht: (2023)
von: Liu, Haotian, et al.
Veröffentlicht: (2023)
Parrot: Multilingual Visual Instruction Tuning
von: Sun, Hai-Long, et al.
Veröffentlicht: (2024)
von: Sun, Hai-Long, et al.
Veröffentlicht: (2024)
Visual Instruction Tuning with Chain of Region-of-Interest
von: Chen, Yixin, et al.
Veröffentlicht: (2025)
von: Chen, Yixin, et al.
Veröffentlicht: (2025)
What Makes Good Instruction-Tuning Data? An In-Context Learning Perspective
von: Han, Guangzeng, et al.
Veröffentlicht: (2026)
von: Han, Guangzeng, et al.
Veröffentlicht: (2026)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
von: He, Junwen, et al.
Veröffentlicht: (2024)
von: He, Junwen, et al.
Veröffentlicht: (2024)
Osprey: Pixel Understanding with Visual Instruction Tuning
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning
von: Xie, Hongxia, et al.
Veröffentlicht: (2024)
von: Xie, Hongxia, et al.
Veröffentlicht: (2024)
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning
von: He, Qingdong, et al.
Veröffentlicht: (2025)
von: He, Qingdong, et al.
Veröffentlicht: (2025)
Bridging Writing Manner Gap in Visual Instruction Tuning by Creating LLM-aligned Instructions
von: Jing, Dong, et al.
Veröffentlicht: (2025)
von: Jing, Dong, et al.
Veröffentlicht: (2025)
Boosting Visual Instruction Tuning with Self-Supervised Guidance
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2026)
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2026)
Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning
von: Safaei, Bardia, et al.
Veröffentlicht: (2025)
von: Safaei, Bardia, et al.
Veröffentlicht: (2025)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Robotic Visual Instruction
von: Li, Yanbang, et al.
Veröffentlicht: (2025)
von: Li, Yanbang, et al.
Veröffentlicht: (2025)
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
von: Zhang, Huanyu, et al.
Veröffentlicht: (2026)
von: Zhang, Huanyu, et al.
Veröffentlicht: (2026)
CoIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity Optimization
von: Yan, Yichen, et al.
Veröffentlicht: (2025)
von: Yan, Yichen, et al.
Veröffentlicht: (2025)
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
von: Ni, Minheng, et al.
Veröffentlicht: (2026)
von: Ni, Minheng, et al.
Veröffentlicht: (2026)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
von: Cui, Hejie, et al.
Veröffentlicht: (2024)
von: Cui, Hejie, et al.
Veröffentlicht: (2024)
How Many Languages Make Good Multilingual Instruction Tuning? A Case Study on BLOOM
von: Ji, Shaoxiong, et al.
Veröffentlicht: (2024)
von: Ji, Shaoxiong, et al.
Veröffentlicht: (2024)
Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
von: Wang, Ziqi, et al.
Veröffentlicht: (2025)
von: Wang, Ziqi, et al.
Veröffentlicht: (2025)
Visually Dehallucinative Instruction Generation
von: Cha, Sungguk, et al.
Veröffentlicht: (2024)
von: Cha, Sungguk, et al.
Veröffentlicht: (2024)
VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
von: Dong, Mingkang, et al.
Veröffentlicht: (2026)
von: Dong, Mingkang, et al.
Veröffentlicht: (2026)
Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2024)
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2024)
VIGC: Visual Instruction Generation and Correction
von: Wang, Bin, et al.
Veröffentlicht: (2023)
von: Wang, Bin, et al.
Veröffentlicht: (2023)
ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
Hierarchical Instruction-aware Embodied Visual Tracking
von: Wu, Kui, et al.
Veröffentlicht: (2025)
von: Wu, Kui, et al.
Veröffentlicht: (2025)
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
von: Hansen, Jacob, et al.
Veröffentlicht: (2025)
von: Hansen, Jacob, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Less is More: High-value Data Selection for Visual Instruction Tuning
von: Liu, Zikang, et al.
Veröffentlicht: (2024) -
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
von: Liu, Zikang, et al.
Veröffentlicht: (2025) -
Reconstructive Visual Instruction Tuning
von: Wang, Haochen, et al.
Veröffentlicht: (2024) -
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
von: Li, Yifan, et al.
Veröffentlicht: (2024) -
EAGLE: Elevating Geometric Reasoning through LLM-empowered Visual Instruction Tuning
von: Li, Zhihao, et al.
Veröffentlicht: (2024)