EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908637017931776 |
|---|---|
| author | Khalid, Umar Munir, Kashif Iqbal, Hasan Farooq, Azib Hua, Jing Rahnavard, Nazanin Chen, Chen Zhu, Victor Ji, Zhengping |
| author_facet | Khalid, Umar Munir, Kashif Iqbal, Hasan Farooq, Azib Hua, Jing Rahnavard, Nazanin Chen, Chen Zhu, Victor Ji, Zhengping |
| contents | Editing complex visual content from ambiguous or partially specified instructions remains a core challenge in vision-language modeling. Existing models can contextualize content but often fail to infer the underlying intent within a reference image or scene, leading to inconsistent or misaligned edits. We introduce the Editing Vision-Language Model (EVLM), a system that interprets ambiguous instructions in conjunction with reference visuals to produce precise, context-aware editing prompts. EVLM's key innovation is a reflective reasoning framework that translates subjective user intent into structured, actionable outputs by aligning with human-rated rationales through Reflection-Aware KL-Divergence Target Optimization (RKTO). By combining Chain-of-Thought (CoT) reasoning with RKTO alignment, EVLM captures fine-grained editing preferences without relying on binary supervision. Trained on a dataset of 30,000 CoT examples with human-annotated rationale quality, EVLM achieves substantial gains in alignment with human intent. Experiments across image, video, 3D, and 4D editing tasks show that EVLM generates coherent and high-quality instructions, providing a scalable foundation for multimodal editing and reasoning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_10566 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing Khalid, Umar Munir, Kashif Iqbal, Hasan Farooq, Azib Hua, Jing Rahnavard, Nazanin Chen, Chen Zhu, Victor Ji, Zhengping Computer Vision and Pattern Recognition Editing complex visual content from ambiguous or partially specified instructions remains a core challenge in vision-language modeling. Existing models can contextualize content but often fail to infer the underlying intent within a reference image or scene, leading to inconsistent or misaligned edits. We introduce the Editing Vision-Language Model (EVLM), a system that interprets ambiguous instructions in conjunction with reference visuals to produce precise, context-aware editing prompts. EVLM's key innovation is a reflective reasoning framework that translates subjective user intent into structured, actionable outputs by aligning with human-rated rationales through Reflection-Aware KL-Divergence Target Optimization (RKTO). By combining Chain-of-Thought (CoT) reasoning with RKTO alignment, EVLM captures fine-grained editing preferences without relying on binary supervision. Trained on a dataset of 30,000 CoT examples with human-annotated rationale quality, EVLM achieves substantial gains in alignment with human intent. Experiments across image, video, 3D, and 4D editing tasks show that EVLM generates coherent and high-quality instructions, providing a scalable foundation for multimodal editing and reasoning. |
| title | EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.10566 |