EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khalid, Umar, Munir, Kashif, Iqbal, Hasan, Farooq, Azib, Hua, Jing, Rahnavard, Nazanin, Chen, Chen, Zhu, Victor, Ji, Zhengping
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908637017931776
author Khalid, Umar
Munir, Kashif
Iqbal, Hasan
Farooq, Azib
Hua, Jing
Rahnavard, Nazanin
Chen, Chen
Zhu, Victor
Ji, Zhengping
author_facet Khalid, Umar
Munir, Kashif
Iqbal, Hasan
Farooq, Azib
Hua, Jing
Rahnavard, Nazanin
Chen, Chen
Zhu, Victor
Ji, Zhengping
contents Editing complex visual content from ambiguous or partially specified instructions remains a core challenge in vision-language modeling. Existing models can contextualize content but often fail to infer the underlying intent within a reference image or scene, leading to inconsistent or misaligned edits. We introduce the Editing Vision-Language Model (EVLM), a system that interprets ambiguous instructions in conjunction with reference visuals to produce precise, context-aware editing prompts. EVLM's key innovation is a reflective reasoning framework that translates subjective user intent into structured, actionable outputs by aligning with human-rated rationales through Reflection-Aware KL-Divergence Target Optimization (RKTO). By combining Chain-of-Thought (CoT) reasoning with RKTO alignment, EVLM captures fine-grained editing preferences without relying on binary supervision. Trained on a dataset of 30,000 CoT examples with human-annotated rationale quality, EVLM achieves substantial gains in alignment with human intent. Experiments across image, video, 3D, and 4D editing tasks show that EVLM generates coherent and high-quality instructions, providing a scalable foundation for multimodal editing and reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10566
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing
Khalid, Umar
Munir, Kashif
Iqbal, Hasan
Farooq, Azib
Hua, Jing
Rahnavard, Nazanin
Chen, Chen
Zhu, Victor
Ji, Zhengping
Computer Vision and Pattern Recognition
Editing complex visual content from ambiguous or partially specified instructions remains a core challenge in vision-language modeling. Existing models can contextualize content but often fail to infer the underlying intent within a reference image or scene, leading to inconsistent or misaligned edits. We introduce the Editing Vision-Language Model (EVLM), a system that interprets ambiguous instructions in conjunction with reference visuals to produce precise, context-aware editing prompts. EVLM's key innovation is a reflective reasoning framework that translates subjective user intent into structured, actionable outputs by aligning with human-rated rationales through Reflection-Aware KL-Divergence Target Optimization (RKTO). By combining Chain-of-Thought (CoT) reasoning with RKTO alignment, EVLM captures fine-grained editing preferences without relying on binary supervision. Trained on a dataset of 30,000 CoT examples with human-annotated rationale quality, EVLM achieves substantial gains in alignment with human intent. Experiments across image, video, 3D, and 4D editing tasks show that EVLM generates coherent and high-quality instructions, providing a scalable foundation for multimodal editing and reasoning.
title EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.10566