Generative Visual Instruction Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hernandez, Jefferson, Villegas, Ruben, Ordonez, Vicente |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2023)
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2023)
Agentic Discovery with Active Hypothesis Exploration for Visual Recognition
von: Koo, Jaywon, et al.
Veröffentlicht: (2026)
von: Koo, Jaywon, et al.
Veröffentlicht: (2026)
GViT: Representing Images as Gaussians for Visual Recognition
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2025)
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2025)
Improving Large Vision and Language Models by Learning from a Panel of Peers
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2025)
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2025)
Personalized Visual Instruction Tuning
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
Comparison Visual Instruction Tuning
von: Lin, Wei, et al.
Veröffentlicht: (2024)
von: Lin, Wei, et al.
Veröffentlicht: (2024)
ProxyThinker: Test-Time Guidance through Small Visual Reasoners
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
von: Koo, Jaywon, et al.
Veröffentlicht: (2025)
von: Koo, Jaywon, et al.
Veröffentlicht: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Osprey: Pixel Understanding with Visual Instruction Tuning
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
Visual Instruction Tuning with Chain of Region-of-Interest
von: Chen, Yixin, et al.
Veröffentlicht: (2025)
von: Chen, Yixin, et al.
Veröffentlicht: (2025)
LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
von: Lai, Bolin, et al.
Veröffentlicht: (2023)
von: Lai, Bolin, et al.
Veröffentlicht: (2023)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
Boosting Visual Instruction Tuning with Self-Supervised Guidance
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2026)
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2026)
MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
von: Zhang, Renrui, et al.
Veröffentlicht: (2024)
von: Zhang, Renrui, et al.
Veröffentlicht: (2024)
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
von: Hansen, Jacob, et al.
Veröffentlicht: (2025)
von: Hansen, Jacob, et al.
Veröffentlicht: (2025)
OASIS: Online Sample Selection for Continual Visual Instruction Tuning
von: Lee, Minjae, et al.
Veröffentlicht: (2025)
von: Lee, Minjae, et al.
Veröffentlicht: (2025)
LLaVA-c: Continual Improved Visual Instruction Tuning
von: Liu, Wenzhuo, et al.
Veröffentlicht: (2025)
von: Liu, Wenzhuo, et al.
Veröffentlicht: (2025)
Visual Generation Tuning
von: Guo, Jiahao, et al.
Veröffentlicht: (2025)
von: Guo, Jiahao, et al.
Veröffentlicht: (2025)
Visually Dehallucinative Instruction Generation
von: Cha, Sungguk, et al.
Veröffentlicht: (2024)
von: Cha, Sungguk, et al.
Veröffentlicht: (2024)
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
von: He, Junwen, et al.
Veröffentlicht: (2024)
von: He, Junwen, et al.
Veröffentlicht: (2024)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
Quality Assessment for AI Generated Images with Instruction Tuning
von: Wang, Jiarui, et al.
Veröffentlicht: (2024)
von: Wang, Jiarui, et al.
Veröffentlicht: (2024)
EAGLE: Elevating Geometric Reasoning through LLM-empowered Visual Instruction Tuning
von: Li, Zhihao, et al.
Veröffentlicht: (2024)
von: Li, Zhihao, et al.
Veröffentlicht: (2024)
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
von: Peng, Wujian, et al.
Veröffentlicht: (2024)
von: Peng, Wujian, et al.
Veröffentlicht: (2024)
Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
von: Wang, Ziqi, et al.
Veröffentlicht: (2025)
von: Wang, Ziqi, et al.
Veröffentlicht: (2025)
Medical Image Understanding Improves Survival Prediction via Visual Instruction Tuning
von: Liu, Xixi, et al.
Veröffentlicht: (2026)
von: Liu, Xixi, et al.
Veröffentlicht: (2026)
ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2023)
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2023)
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
PropTest: Automatic Property Testing for Improved Visual Programming
von: Koo, Jaywon, et al.
Veröffentlicht: (2024)
von: Koo, Jaywon, et al.
Veröffentlicht: (2024)
Less is More: High-value Data Selection for Visual Instruction Tuning
von: Liu, Zikang, et al.
Veröffentlicht: (2024)
von: Liu, Zikang, et al.
Veröffentlicht: (2024)
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning
von: Vogel, Alexander, et al.
Veröffentlicht: (2025)
von: Vogel, Alexander, et al.
Veröffentlicht: (2025)
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
von: li, Bonan, et al.
Veröffentlicht: (2025)
von: li, Bonan, et al.
Veröffentlicht: (2025)
Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction Tuning
von: Gou, Yunhao, et al.
Veröffentlicht: (2025)
von: Gou, Yunhao, et al.
Veröffentlicht: (2025)
Reconstructive Visual Instruction Tuning
von: Wang, Haochen, et al.
Veröffentlicht: (2024)
von: Wang, Haochen, et al.
Veröffentlicht: (2024)
Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning
von: Safaei, Bardia, et al.
Veröffentlicht: (2025)
von: Safaei, Bardia, et al.
Veröffentlicht: (2025)
Streaming Video Instruction Tuning
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2023) -
Agentic Discovery with Active Hypothesis Exploration for Visual Recognition
von: Koo, Jaywon, et al.
Veröffentlicht: (2026) -
GViT: Representing Images as Gaussians for Visual Recognition
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2025) -
Improving Large Vision and Language Models by Learning from a Panel of Peers
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2025) -
Personalized Visual Instruction Tuning
von: Pi, Renjie, et al.
Veröffentlicht: (2024)