Do-Undo Bench: Reversibility for Action Understanding in Image Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mahajan, Shweta, Kadambi, Shreya, Le, Hoang, Yasarla, Rajeev, Bhattacharyya, Apratim, Hayat, Munawar, Porikli, Fatih
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910220226134016
author Mahajan, Shweta
Kadambi, Shreya
Le, Hoang
Yasarla, Rajeev
Bhattacharyya, Apratim
Hayat, Munawar
Porikli, Fatih
author_facet Mahajan, Shweta
Kadambi, Shreya
Le, Hoang
Yasarla, Rajeev
Bhattacharyya, Apratim
Hayat, Munawar
Porikli, Fatih
contents We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transformations driven by real-world actions. Unlike prior work that relies on prompt-based image generation and editing to perform action-conditioned image manipulation, our training hypothesis requires models to simulate the outcome of a real-world action and then reverse it to the original state. This forward-reverse requirement tests genuine cause-and-effect understanding rather than stylistic or semantic edits. We curate a high-quality benchmark of reversible actions from real-world scenarios to enable robust action grounding. Our experiments reveal that current models struggle with action reversibility, highlighting the need to evaluate action understanding. Do-Undo provides an intuitive testbed for evaluating and advancing action-aware generation in multimodal systems that must reason about real-world dynamics.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13609
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do-Undo Bench: Reversibility for Action Understanding in Image Generation
Mahajan, Shweta
Kadambi, Shreya
Le, Hoang
Yasarla, Rajeev
Bhattacharyya, Apratim
Hayat, Munawar
Porikli, Fatih
Computer Vision and Pattern Recognition
Machine Learning
We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transformations driven by real-world actions. Unlike prior work that relies on prompt-based image generation and editing to perform action-conditioned image manipulation, our training hypothesis requires models to simulate the outcome of a real-world action and then reverse it to the original state. This forward-reverse requirement tests genuine cause-and-effect understanding rather than stylistic or semantic edits. We curate a high-quality benchmark of reversible actions from real-world scenarios to enable robust action grounding. Our experiments reveal that current models struggle with action reversibility, highlighting the need to evaluate action understanding. Do-Undo provides an intuitive testbed for evaluating and advancing action-aware generation in multimodal systems that must reason about real-world dynamics.
title Do-Undo Bench: Reversibility for Action Understanding in Image Generation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2512.13609