How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Huanyu, Bai, Xuehai, Li, Chengzu, Liang, Chen, Tian, Haochen, Li, Haodong, An, Ruichuan, Zhang, Yifan, Korhonen, Anna, Zhang, Zhang, Wang, Liang, Tan, Tieniu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VIBE: Visual Instruction Based Editor
von: Alekseenko, Grigorii, et al.
Veröffentlicht: (2026)
von: Alekseenko, Grigorii, et al.
Veröffentlicht: (2026)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
von: Bai, Xuehai, et al.
Veröffentlicht: (2026)
von: Bai, Xuehai, et al.
Veröffentlicht: (2026)
Reconstructive Visual Instruction Tuning
von: Wang, Haochen, et al.
Veröffentlicht: (2024)
von: Wang, Haochen, et al.
Veröffentlicht: (2024)
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
Visual Planning: Let's Think Only with Images
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
IHEval: Evaluating Language Models on Following the Instruction Hierarchy
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
Semantic Map-based Generation of Navigation Instructions
von: Li, Chengzu, et al.
Veröffentlicht: (2024)
von: Li, Chengzu, et al.
Veröffentlicht: (2024)
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following
von: Li, Shufan, et al.
Veröffentlicht: (2023)
von: Li, Shufan, et al.
Veröffentlicht: (2023)
Personalized Visual Instruction Tuning
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
von: Cui, Hejie, et al.
Veröffentlicht: (2024)
von: Cui, Hejie, et al.
Veröffentlicht: (2024)
ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions
von: Yang, Donglu, et al.
Veröffentlicht: (2025)
von: Yang, Donglu, et al.
Veröffentlicht: (2025)
Robotic Visual Instruction
von: Li, Yanbang, et al.
Veröffentlicht: (2025)
von: Li, Yanbang, et al.
Veröffentlicht: (2025)
A Systematic Examination of Preference Learning through the Lens of Instruction-Following
von: Kim, Joongwon, et al.
Veröffentlicht: (2024)
von: Kim, Joongwon, et al.
Veröffentlicht: (2024)
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
von: Ni, Minheng, et al.
Veröffentlicht: (2026)
von: Ni, Minheng, et al.
Veröffentlicht: (2026)
MathChat: Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Interactions
von: Liang, Zhenwen, et al.
Veröffentlicht: (2024)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2024)
TimeRAF: Retrieval-Augmented Foundation model for Zero-shot Time Series Forecasting
von: Zhang, Huanyu, et al.
Veröffentlicht: (2024)
von: Zhang, Huanyu, et al.
Veröffentlicht: (2024)
HIVE: Harnessing Human Feedback for Instructional Visual Editing
von: Zhang, Shu, et al.
Veröffentlicht: (2023)
von: Zhang, Shu, et al.
Veröffentlicht: (2023)
Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning
von: He, Qingdong, et al.
Veröffentlicht: (2025)
von: He, Qingdong, et al.
Veröffentlicht: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Visual Autoregressive Modeling for Instruction-Guided Image Editing
von: Mao, Qingyang, et al.
Veröffentlicht: (2025)
von: Mao, Qingyang, et al.
Veröffentlicht: (2025)
Automatic Layout Planning for Visually-Rich Documents with Instruction-Following Models
von: Zhu, Wanrong, et al.
Veröffentlicht: (2024)
von: Zhu, Wanrong, et al.
Veröffentlicht: (2024)
EditWorld: Simulating World Dynamics for Instruction-Following Image Editing
von: Yang, Ling, et al.
Veröffentlicht: (2024)
von: Yang, Ling, et al.
Veröffentlicht: (2024)
Human Image Generation: A Comprehensive Survey
von: Jia, Zhen, et al.
Veröffentlicht: (2022)
von: Jia, Zhen, et al.
Veröffentlicht: (2022)
Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2025)
CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models
von: LI, Yizhi, et al.
Veröffentlicht: (2024)
von: LI, Yizhi, et al.
Veröffentlicht: (2024)
XIFBench: Evaluating Large Language Models on Multilingual Instruction Following
von: Li, Zhenyu, et al.
Veröffentlicht: (2025)
von: Li, Zhenyu, et al.
Veröffentlicht: (2025)
Instruction Following without Instruction Tuning
von: Hewitt, John, et al.
Veröffentlicht: (2024)
von: Hewitt, John, et al.
Veröffentlicht: (2024)
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
von: Chen, Zhihong, et al.
Veröffentlicht: (2025)
von: Chen, Zhihong, et al.
Veröffentlicht: (2025)
LERa: Replanning with Visual Feedback in Instruction Following
von: Pchelintsev, Svyatoslav, et al.
Veröffentlicht: (2025)
von: Pchelintsev, Svyatoslav, et al.
Veröffentlicht: (2025)
PEARL: Personalized Streaming Video Understanding Model
von: Zheng, Yuanhong, et al.
Veröffentlicht: (2026)
von: Zheng, Yuanhong, et al.
Veröffentlicht: (2026)
Learning to Instruct for Visual Instruction Tuning
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
VIGC: Visual Instruction Generation and Correction
von: Wang, Bin, et al.
Veröffentlicht: (2023)
von: Wang, Bin, et al.
Veröffentlicht: (2023)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies
von: Wang, Chenglin, et al.
Veröffentlicht: (2025)
von: Wang, Chenglin, et al.
Veröffentlicht: (2025)
Can Language Models Follow Multiple Turns of Entangled Instructions?
von: Han, Chi, et al.
Veröffentlicht: (2025)
von: Han, Chi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VIBE: Visual Instruction Based Editor
von: Alekseenko, Grigorii, et al.
Veröffentlicht: (2026) -
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025) -
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
von: Li, Chengzu, et al.
Veröffentlicht: (2026) -
Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025) -
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
von: Bai, Xuehai, et al.
Veröffentlicht: (2026)