Less is More: High-value Data Selection for Visual Instruction Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zikang, Zhou, Kun, Zhao, Wayne Xin, Gao, Dawei, Li, Yaliang, Wen, Ji-Rong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
von: Li, Yifan, et al.
Veröffentlicht: (2024)
von: Li, Yifan, et al.
Veröffentlicht: (2024)
PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
von: Li, Yifan, et al.
Veröffentlicht: (2025)
von: Li, Yifan, et al.
Veröffentlicht: (2025)
Less Is More? Selective Visual Attention to High-Importance Regions for Multimodal Radiology Summarization
von: Naznin, Mst. Fahmida Sultana, et al.
Veröffentlicht: (2026)
von: Naznin, Mst. Fahmida Sultana, et al.
Veröffentlicht: (2026)
Exploring the Design Space of Visual Context Representation in Video MLLMs
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
Data-CUBE: Data Curriculum for Instruction-based Sentence Representation Learning
von: Min, Yingqian, et al.
Veröffentlicht: (2024)
von: Min, Yingqian, et al.
Veröffentlicht: (2024)
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
von: You, Zebin, et al.
Veröffentlicht: (2025)
von: You, Zebin, et al.
Veröffentlicht: (2025)
Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2024)
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2024)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2024)
von: Zhou, Shijie, et al.
Veröffentlicht: (2024)
Towards Event-oriented Long Video Understanding
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking
von: Cheng, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Cheng, Xiaoxue, et al.
Veröffentlicht: (2025)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
Reconstructive Visual Instruction Tuning
von: Wang, Haochen, et al.
Veröffentlicht: (2024)
von: Wang, Haochen, et al.
Veröffentlicht: (2024)
Parrot: Multilingual Visual Instruction Tuning
von: Sun, Hai-Long, et al.
Veröffentlicht: (2024)
von: Sun, Hai-Long, et al.
Veröffentlicht: (2024)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
Learning to Instruct for Visual Instruction Tuning
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
ReasoningLM: Enabling Structural Subgraph Reasoning in Pre-trained Language Models for Question Answering over Knowledge Graph
von: Jiang, Jinhao, et al.
Veröffentlicht: (2023)
von: Jiang, Jinhao, et al.
Veröffentlicht: (2023)
MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning
von: Ma, Yiwei, et al.
Veröffentlicht: (2025)
von: Ma, Yiwei, et al.
Veröffentlicht: (2025)
Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective
von: Yue, Zihao, et al.
Veröffentlicht: (2024)
von: Yue, Zihao, et al.
Veröffentlicht: (2024)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
Improved Baselines with Visual Instruction Tuning
von: Liu, Haotian, et al.
Veröffentlicht: (2023)
von: Liu, Haotian, et al.
Veröffentlicht: (2023)
More Text, Less Point: Towards 3D Data-Efficient Point-Language Understanding
von: Tang, Yuan, et al.
Veröffentlicht: (2024)
von: Tang, Yuan, et al.
Veröffentlicht: (2024)
MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
von: Xu, Zhe, et al.
Veröffentlicht: (2025)
von: Xu, Zhe, et al.
Veröffentlicht: (2025)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights
von: Wen, Xin, et al.
Veröffentlicht: (2024)
von: Wen, Xin, et al.
Veröffentlicht: (2024)
Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability
von: Hao, Tianxiang, et al.
Veröffentlicht: (2023)
von: Hao, Tianxiang, et al.
Veröffentlicht: (2023)
MoExtend: Tuning New Experts for Modality and Task Extension
von: Zhong, Shanshan, et al.
Veröffentlicht: (2024)
von: Zhong, Shanshan, et al.
Veröffentlicht: (2024)
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
von: Bi, Jinhe, et al.
Veröffentlicht: (2024)
von: Bi, Jinhe, et al.
Veröffentlicht: (2024)
Unleashing the Power of Data Tsunami: A Comprehensive Survey on Data Assessment and Selection for Instruction Tuning of Language Models
von: Qin, Yulei, et al.
Veröffentlicht: (2024)
von: Qin, Yulei, et al.
Veröffentlicht: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
von: Li, Bin, et al.
Veröffentlicht: (2025)
von: Li, Bin, et al.
Veröffentlicht: (2025)
Getting More Juice Out of Your Data: Hard Pair Refinement Enhances Visual-Language Models Without Extra Data
von: Wang, Haonan, et al.
Veröffentlicht: (2023)
von: Wang, Haonan, et al.
Veröffentlicht: (2023)
Otter: A Multi-Modal Model with In-Context Instruction Tuning
von: Li, Bo, et al.
Veröffentlicht: (2023)
von: Li, Bo, et al.
Veröffentlicht: (2023)
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
von: Li, Yifan, et al.
Veröffentlicht: (2025)
von: Li, Yifan, et al.
Veröffentlicht: (2025)
Little Data, Big Impact: Privacy-Aware Visual Language Models via Minimal Tuning
von: Samson, Laurens, et al.
Veröffentlicht: (2024)
von: Samson, Laurens, et al.
Veröffentlicht: (2024)
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
von: Wang, Yanan, et al.
Veröffentlicht: (2025)
von: Wang, Yanan, et al.
Veröffentlicht: (2025)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
von: Guo, Jarvis, et al.
Veröffentlicht: (2024)
von: Guo, Jarvis, et al.
Veröffentlicht: (2024)
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
von: Cocchi, Federico, et al.
Veröffentlicht: (2025)
von: Cocchi, Federico, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
von: Liu, Zikang, et al.
Veröffentlicht: (2025) -
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023) -
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
von: Li, Yifan, et al.
Veröffentlicht: (2024) -
PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
von: Liu, Zikang, et al.
Veröffentlicht: (2025) -
Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
von: Li, Yifan, et al.
Veröffentlicht: (2025)