Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Jing, Zhang, Zhaoyang, Shen, Yantao, Cai, Jiarui, Yang, Shuo, Wu, Jiajun, Xia, Wei, Tu, Zhuowen, Soatto, Stefano |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcement-aware Knowledge Distillation for LLM Reasoning
by: Zhang, Zhaoyang, et al.
Published: (2026)
by: Zhang, Zhaoyang, et al.
Published: (2026)
ELODI: Ensemble Logit Difference Inhibition for Positive-Congruent Training
by: Zhao, Yue, et al.
Published: (2022)
by: Zhao, Yue, et al.
Published: (2022)
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
by: Zhang, Zhaoyang, et al.
Published: (2023)
by: Zhang, Zhaoyang, et al.
Published: (2023)
Non-autoregressive Sequence-to-Sequence Vision-Language Models
by: Shi, Kunyu, et al.
Published: (2024)
by: Shi, Kunyu, et al.
Published: (2024)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
by: Ram, Dhananjay, et al.
Published: (2025)
by: Ram, Dhananjay, et al.
Published: (2025)
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models
by: Liu, Xiaoze, et al.
Published: (2026)
by: Liu, Xiaoze, et al.
Published: (2026)
Asymmetric Actor-Critic for Multi-turn LLM Agents
by: Jiang, Shuli, et al.
Published: (2026)
by: Jiang, Shuli, et al.
Published: (2026)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers
by: Srivastava, Divyansh, et al.
Published: (2025)
by: Srivastava, Divyansh, et al.
Published: (2025)
On the Scalability of Diffusion-based Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Point'n Move: Interactive Scene Object Manipulation on Gaussian Splatting Radiance Fields
by: Huang, Jiajun, et al.
Published: (2023)
by: Huang, Jiajun, et al.
Published: (2023)
InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
by: Cai, Xinhao, et al.
Published: (2025)
by: Cai, Xinhao, et al.
Published: (2025)
Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model
by: Li, Xiaolong, et al.
Published: (2024)
by: Li, Xiaolong, et al.
Published: (2024)
InstructOCR: Instruction Boosting Scene Text Spotting
by: Duan, Chen, et al.
Published: (2024)
by: Duan, Chen, et al.
Published: (2024)
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
by: Zeng, Guanning, et al.
Published: (2025)
by: Zeng, Guanning, et al.
Published: (2025)
Tangent Transformers for Composition, Privacy and Removal
by: Liu, Tian Yu, et al.
Published: (2023)
by: Liu, Tian Yu, et al.
Published: (2023)
AI Agents as Universal Task Solvers
by: Achille, Alessandro, et al.
Published: (2025)
by: Achille, Alessandro, et al.
Published: (2025)
Cycles of Thought: Measuring LLM Confidence through Stable Explanations
by: Becker, Evan, et al.
Published: (2024)
by: Becker, Evan, et al.
Published: (2024)
Robust Planning for Autonomous Driving via Mixed Adversarial Diffusion Predictions
by: Zhao, Albert, et al.
Published: (2025)
by: Zhao, Albert, et al.
Published: (2025)
EvoMAS: Evolutionary Generation of Multi-Agent Systems
by: Hu, Yuntong, et al.
Published: (2026)
by: Hu, Yuntong, et al.
Published: (2026)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
Mixed-Query Transformer: A Unified Image Segmentation Architecture
by: Wang, Pei, et al.
Published: (2024)
by: Wang, Pei, et al.
Published: (2024)
TransBridge: Boost 3D Object Detection by Scene-Level Completion with Transformer Decoder
by: Meng, Qinghao, et al.
Published: (2025)
by: Meng, Qinghao, et al.
Published: (2025)
IPCGRL: Language-Instructed Reinforcement Learning for Procedural Level Generation
by: Baek, In-Chang, et al.
Published: (2025)
by: Baek, In-Chang, et al.
Published: (2025)
Open-World Dynamic Prompt and Continual Visual Representation Learning
by: Kim, Youngeun, et al.
Published: (2024)
by: Kim, Youngeun, et al.
Published: (2024)
e1: Learning Adaptive Control of Reasoning Effort
by: Kleinman, Michael, et al.
Published: (2025)
by: Kleinman, Michael, et al.
Published: (2025)
Restoration by Generation with Constrained Priors
by: Ding, Zheng, et al.
Published: (2023)
by: Ding, Zheng, et al.
Published: (2023)
Enhancing Vision-Language Pre-training with Rich Supervisions
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots
by: Lao, Dong, et al.
Published: (2023)
by: Lao, Dong, et al.
Published: (2023)
Conjuring Semantic Similarity
by: Liu, Tian Yu, et al.
Published: (2024)
by: Liu, Tian Yu, et al.
Published: (2024)
From Scene to Object: Text-Guided Dual-Gaze Prediction
by: Ke, Zehong, et al.
Published: (2026)
by: Ke, Zehong, et al.
Published: (2026)
TokenCompose: Text-to-Image Diffusion with Token-level Supervision
by: Wang, Zirui, et al.
Published: (2023)
by: Wang, Zirui, et al.
Published: (2023)
PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
by: Zhang, Xiang, et al.
Published: (2026)
by: Zhang, Xiang, et al.
Published: (2026)
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
by: Zancato, Luca, et al.
Published: (2024)
by: Zancato, Luca, et al.
Published: (2024)
InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
by: Yu, Haoran, et al.
Published: (2025)
by: Yu, Haoran, et al.
Published: (2025)
SirenPose: Dynamic Scene Reconstruction via Geometric Supervision
by: Cai, Kaitong, et al.
Published: (2025)
by: Cai, Kaitong, et al.
Published: (2025)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
by: Ren, Yong, et al.
Published: (2026)
by: Ren, Yong, et al.
Published: (2026)
Critical Learning Periods Emerge Even in Deep Linear Networks
by: Kleinman, Michael, et al.
Published: (2023)
by: Kleinman, Michael, et al.
Published: (2023)
Priming: Hybrid State Space Models From Pre-trained Transformers
by: Chattopadhyay, Aditya, et al.
Published: (2026)
by: Chattopadhyay, Aditya, et al.
Published: (2026)
TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction
by: Huang, Yiyao, et al.
Published: (2025)
by: Huang, Yiyao, et al.
Published: (2025)
Similar Items
-
Reinforcement-aware Knowledge Distillation for LLM Reasoning
by: Zhang, Zhaoyang, et al.
Published: (2026) -
ELODI: Ensemble Logit Difference Inhibition for Positive-Congruent Training
by: Zhao, Yue, et al.
Published: (2022) -
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
by: Zhang, Zhaoyang, et al.
Published: (2023) -
Non-autoregressive Sequence-to-Sequence Vision-Language Models
by: Shi, Kunyu, et al.
Published: (2024) -
Learning to Focus: Focal Attention for Selective and Scalable Transformers
by: Ram, Dhananjay, et al.
Published: (2025)