Learning to Instruct for Visual Instruction Tuning
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhou, Zhihan, Hong, Feng, Luo, Jiaan, Yao, Jiangchao, Li, Dongsheng, Han, Bo, Zhang, Ya, Wang, Yanfeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Decouple before Align: Visual Disentanglement Enhances Prompt Tuning
por: Zhang, Fei, et al.
Publicado: (2025)
por: Zhang, Fei, et al.
Publicado: (2025)
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
por: Jia, Yiming, et al.
Publicado: (2025)
por: Jia, Yiming, et al.
Publicado: (2025)
Parrot: Multilingual Visual Instruction Tuning
por: Sun, Hai-Long, et al.
Publicado: (2024)
por: Sun, Hai-Long, et al.
Publicado: (2024)
SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
por: Mao, Zhenjie, et al.
Publicado: (2025)
por: Mao, Zhenjie, et al.
Publicado: (2025)
Instruct-Imagen: Image Generation with Multi-modal Instruction
por: Hu, Hexiang, et al.
Publicado: (2024)
por: Hu, Hexiang, et al.
Publicado: (2024)
MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
por: Liu, Jihao, et al.
Publicado: (2024)
por: Liu, Jihao, et al.
Publicado: (2024)
G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models
por: Zhang, Tianjiao, et al.
Publicado: (2025)
por: Zhang, Tianjiao, et al.
Publicado: (2025)
Reconstructive Visual Instruction Tuning
por: Wang, Haochen, et al.
Publicado: (2024)
por: Wang, Haochen, et al.
Publicado: (2024)
Improved Baselines with Visual Instruction Tuning
por: Liu, Haotian, et al.
Publicado: (2023)
por: Liu, Haotian, et al.
Publicado: (2023)
Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection
por: Yu, Geng, et al.
Publicado: (2024)
por: Yu, Geng, et al.
Publicado: (2024)
ReMamber: Referring Image Segmentation with Mamba Twister
por: Yang, Yuhuan, et al.
Publicado: (2024)
por: Yang, Yuhuan, et al.
Publicado: (2024)
InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
por: Long, Yuxing, et al.
Publicado: (2024)
por: Long, Yuxing, et al.
Publicado: (2024)
M^3Builder: A Multi-Agent System for Automated Machine Learning in Medical Imaging
por: Feng, Jinghao, et al.
Publicado: (2025)
por: Feng, Jinghao, et al.
Publicado: (2025)
Dual-granularity Sinkhorn Distillation for Enhanced Learning from Long-tailed Noisy Data
por: Hong, Feng, et al.
Publicado: (2025)
por: Hong, Feng, et al.
Publicado: (2025)
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
por: Zou, Bo, et al.
Publicado: (2024)
por: Zou, Bo, et al.
Publicado: (2024)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
por: Wang, Haochen, et al.
Publicado: (2025)
por: Wang, Haochen, et al.
Publicado: (2025)
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
por: Cocchi, Federico, et al.
Publicado: (2025)
por: Cocchi, Federico, et al.
Publicado: (2025)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
por: Seyfioglu, Mehmet Saygin, et al.
Publicado: (2023)
por: Seyfioglu, Mehmet Saygin, et al.
Publicado: (2023)
MANTIS: Interleaved Multi-Image Instruction Tuning
por: Jiang, Dongfu, et al.
Publicado: (2024)
por: Jiang, Dongfu, et al.
Publicado: (2024)
EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning
por: Xie, Hongxia, et al.
Publicado: (2024)
por: Xie, Hongxia, et al.
Publicado: (2024)
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
por: Yan, Qiao, et al.
Publicado: (2025)
por: Yan, Qiao, et al.
Publicado: (2025)
Dr-LLaVA: Visual Instruction Tuning with Symbolic Clinical Grounding
por: Sun, Shenghuan, et al.
Publicado: (2024)
por: Sun, Shenghuan, et al.
Publicado: (2024)
InstructOCR: Instruction Boosting Scene Text Spotting
por: Duan, Chen, et al.
Publicado: (2024)
por: Duan, Chen, et al.
Publicado: (2024)
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants
por: Bansal, Hritik, et al.
Publicado: (2024)
por: Bansal, Hritik, et al.
Publicado: (2024)
InstructDET: Diversifying Referring Object Detection with Generalized Instructions
por: Dang, Ronghao, et al.
Publicado: (2023)
por: Dang, Ronghao, et al.
Publicado: (2023)
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
por: Fan, Ziqing, et al.
Publicado: (2025)
por: Fan, Ziqing, et al.
Publicado: (2025)
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
por: Yao, Louie Hong, et al.
Publicado: (2025)
por: Yao, Louie Hong, et al.
Publicado: (2025)
Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models
por: Li, Shengzhi, et al.
Publicado: (2024)
por: Li, Shengzhi, et al.
Publicado: (2024)
LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
por: Fujitake, Masato
Publicado: (2024)
por: Fujitake, Masato
Publicado: (2024)
ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild
por: Masry, Ahmed, et al.
Publicado: (2024)
por: Masry, Ahmed, et al.
Publicado: (2024)
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
por: Luo, Run, et al.
Publicado: (2025)
por: Luo, Run, et al.
Publicado: (2025)
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
InstructEdit: Instruction-based Knowledge Editing for Large Language Models
por: Zhang, Ningyu, et al.
Publicado: (2024)
por: Zhang, Ningyu, et al.
Publicado: (2024)
Automatic Layout Planning for Visually-Rich Documents with Instruction-Following Models
por: Zhu, Wanrong, et al.
Publicado: (2024)
por: Zhu, Wanrong, et al.
Publicado: (2024)
Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video
por: Li, Bin, et al.
Publicado: (2022)
por: Li, Bin, et al.
Publicado: (2022)
KARL: Knowledge-Aware Reasoning and Reinforcement Learning for Knowledge-Intensive Visual Grounding
por: Ma, Xinyu, et al.
Publicado: (2025)
por: Ma, Xinyu, et al.
Publicado: (2025)
OmniGen2: Towards Instruction-Aligned Multimodal Generation
por: Wu, Chenyuan, et al.
Publicado: (2025)
por: Wu, Chenyuan, et al.
Publicado: (2025)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
por: Yu, Shoubin, et al.
Publicado: (2025)
por: Yu, Shoubin, et al.
Publicado: (2025)
VGR: Visual Grounded Reasoning
por: Wang, Jiacong, et al.
Publicado: (2025)
por: Wang, Jiacong, et al.
Publicado: (2025)
Cost-effective Instruction Learning for Pathology Vision and Language Analysis
por: Chen, Kaitao, et al.
Publicado: (2024)
por: Chen, Kaitao, et al.
Publicado: (2024)
Ejemplares similares
-
Decouple before Align: Visual Disentanglement Enhances Prompt Tuning
por: Zhang, Fei, et al.
Publicado: (2025) -
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
por: Jia, Yiming, et al.
Publicado: (2025) -
Parrot: Multilingual Visual Instruction Tuning
por: Sun, Hai-Long, et al.
Publicado: (2024) -
SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
por: Mao, Zhenjie, et al.
Publicado: (2025) -
Instruct-Imagen: Image Generation with Multi-modal Instruction
por: Hu, Hexiang, et al.
Publicado: (2024)