Do-Undo Bench: Reversibility for Action Understanding in Image Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910220226134016 |
|---|---|
| author | Mahajan, Shweta Kadambi, Shreya Le, Hoang Yasarla, Rajeev Bhattacharyya, Apratim Hayat, Munawar Porikli, Fatih |
| author_facet | Mahajan, Shweta Kadambi, Shreya Le, Hoang Yasarla, Rajeev Bhattacharyya, Apratim Hayat, Munawar Porikli, Fatih |
| contents | We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transformations driven by real-world actions. Unlike prior work that relies on prompt-based image generation and editing to perform action-conditioned image manipulation, our training hypothesis requires models to simulate the outcome of a real-world action and then reverse it to the original state. This forward-reverse requirement tests genuine cause-and-effect understanding rather than stylistic or semantic edits. We curate a high-quality benchmark of reversible actions from real-world scenarios to enable robust action grounding. Our experiments reveal that current models struggle with action reversibility, highlighting the need to evaluate action understanding. Do-Undo provides an intuitive testbed for evaluating and advancing action-aware generation in multimodal systems that must reason about real-world dynamics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_13609 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Do-Undo Bench: Reversibility for Action Understanding in Image Generation Mahajan, Shweta Kadambi, Shreya Le, Hoang Yasarla, Rajeev Bhattacharyya, Apratim Hayat, Munawar Porikli, Fatih Computer Vision and Pattern Recognition Machine Learning We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transformations driven by real-world actions. Unlike prior work that relies on prompt-based image generation and editing to perform action-conditioned image manipulation, our training hypothesis requires models to simulate the outcome of a real-world action and then reverse it to the original state. This forward-reverse requirement tests genuine cause-and-effect understanding rather than stylistic or semantic edits. We curate a high-quality benchmark of reversible actions from real-world scenarios to enable robust action grounding. Our experiments reveal that current models struggle with action reversibility, highlighting the need to evaluate action understanding. Do-Undo provides an intuitive testbed for evaluating and advancing action-aware generation in multimodal systems that must reason about real-world dynamics. |
| title | Do-Undo Bench: Reversibility for Action Understanding in Image Generation |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2512.13609 |