Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tan, Jing, Zhang, Zhaoyang, Shen, Yantao, Cai, Jiarui, Yang, Shuo, Wu, Jiajun, Xia, Wei, Tu, Zhuowen, Soatto, Stefano
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911359921291264
author Tan, Jing
Zhang, Zhaoyang
Shen, Yantao
Cai, Jiarui
Yang, Shuo
Wu, Jiajun
Xia, Wei
Tu, Zhuowen
Soatto, Stefano
author_facet Tan, Jing
Zhang, Zhaoyang
Shen, Yantao
Cai, Jiarui
Yang, Shuo
Wu, Jiajun
Xia, Wei
Tu, Zhuowen
Soatto, Stefano
contents We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects in a scene through natural language poses a challenge for multimodal generation systems. While existing text-based manipulation methods can adjust appearance or style, they struggle to perform object-level geometric transformations-such as translating, rotating, or resizing objects-due to scarce paired supervision and pixel-level optimization limits. Talk2Move employs Group Relative Policy Optimization (GRPO) to explore geometric actions through diverse rollouts generated from input images and lightweight textual variations, removing the need for costly paired data. A spatial reward guided model aligns geometric transformations with linguistic description, while off-policy step evaluation and active step sampling improve learning efficiency by focusing on informative transformation stages. Furthermore, we design object-centric spatial rewards that evaluate displacement, rotation, and scaling behaviors directly, enabling interpretable and coherent transformations. Experiments on curated benchmarks demonstrate that Talk2Move achieves precise, consistent, and semantically faithful object transformations, outperforming existing text-guided editing approaches in both spatial accuracy and scene coherence.
format Preprint
id arxiv_https___arxiv_org_abs_2601_02356
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
Tan, Jing
Zhang, Zhaoyang
Shen, Yantao
Cai, Jiarui
Yang, Shuo
Wu, Jiajun
Xia, Wei
Tu, Zhuowen
Soatto, Stefano
Computer Vision and Pattern Recognition
We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects in a scene through natural language poses a challenge for multimodal generation systems. While existing text-based manipulation methods can adjust appearance or style, they struggle to perform object-level geometric transformations-such as translating, rotating, or resizing objects-due to scarce paired supervision and pixel-level optimization limits. Talk2Move employs Group Relative Policy Optimization (GRPO) to explore geometric actions through diverse rollouts generated from input images and lightweight textual variations, removing the need for costly paired data. A spatial reward guided model aligns geometric transformations with linguistic description, while off-policy step evaluation and active step sampling improve learning efficiency by focusing on informative transformation stages. Furthermore, we design object-centric spatial rewards that evaluate displacement, rotation, and scaling behaviors directly, enabling interpretable and coherent transformations. Experiments on curated benchmarks demonstrate that Talk2Move achieves precise, consistent, and semantically faithful object transformations, outperforming existing text-guided editing approaches in both spatial accuracy and scene coherence.
title Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.02356