Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
Fuente:
arXiv
Guardado en:
| Autores principales: | Gu, Bohai, Wu, Taiyi, Du, Dazhao, Liu, Jian, Yang, Shuai, Zhao, Xiaotong, Zhao, Alan, Guo, Song |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
por: Gu, Bohai, et al.
Publicado: (2026)
por: Gu, Bohai, et al.
Publicado: (2026)
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
por: Du, Dazhao, et al.
Publicado: (2026)
por: Du, Dazhao, et al.
Publicado: (2026)
Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
por: Zhao, Lingxiao, et al.
Publicado: (2025)
por: Zhao, Lingxiao, et al.
Publicado: (2025)
ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation
por: Yang, Songlin, et al.
Publicado: (2026)
por: Yang, Songlin, et al.
Publicado: (2026)
Coherent Video Inpainting Using Optical Flow-Guided Efficient Diffusion
por: Gu, Bohai, et al.
Publicado: (2024)
por: Gu, Bohai, et al.
Publicado: (2024)
TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation
por: Zhong, Linqing, et al.
Publicado: (2024)
por: Zhong, Linqing, et al.
Publicado: (2024)
Predicting the Future by Retrieving the Past
por: Du, Dazhao, et al.
Publicado: (2025)
por: Du, Dazhao, et al.
Publicado: (2025)
From Visual Synthesis to Interactive Worlds: Toward Production-Ready 3D Asset Generation
por: Wu, Jiafeng, et al.
Publicado: (2026)
por: Wu, Jiafeng, et al.
Publicado: (2026)
Flow-Guided Diffusion for Video Inpainting
por: Gu, Bohai, et al.
Publicado: (2023)
por: Gu, Bohai, et al.
Publicado: (2023)
Online Reasoning Video Object Segmentation
por: Liu, Jinyuan, et al.
Publicado: (2026)
por: Liu, Jinyuan, et al.
Publicado: (2026)
CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition
por: Yang, Hongji, et al.
Publicado: (2026)
por: Yang, Hongji, et al.
Publicado: (2026)
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
por: Li, Yifan, et al.
Publicado: (2025)
por: Li, Yifan, et al.
Publicado: (2025)
VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control
por: Tu, Yuanpeng, et al.
Publicado: (2025)
por: Tu, Yuanpeng, et al.
Publicado: (2025)
NYC-Event-VPR: A Large-Scale High-Resolution Event-Based Visual Place Recognition Dataset in Dense Urban Environments
por: Pan, Taiyi, et al.
Publicado: (2024)
por: Pan, Taiyi, et al.
Publicado: (2024)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
por: Tzachor, Issar, et al.
Publicado: (2026)
por: Tzachor, Issar, et al.
Publicado: (2026)
Anything in Any Scene: Photorealistic Video Object Insertion
por: Bai, Chen, et al.
Publicado: (2024)
por: Bai, Chen, et al.
Publicado: (2024)
Efficient Motion-Aware Video MLLM
por: Zhao, Zijia, et al.
Publicado: (2025)
por: Zhao, Zijia, et al.
Publicado: (2025)
DreamInsert: Zero-Shot Image-to-Video Object Insertion from A Single Image
por: Zhao, Qi, et al.
Publicado: (2025)
por: Zhao, Qi, et al.
Publicado: (2025)
Disruptions as Opportunities
por: Sun, Taiyi
Publicado: (2023)
por: Sun, Taiyi
Publicado: (2023)
VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning
por: Xu, Zishan, et al.
Publicado: (2025)
por: Xu, Zishan, et al.
Publicado: (2025)
The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
por: Gao, Shibo, et al.
Publicado: (2025)
por: Gao, Shibo, et al.
Publicado: (2025)
GO-Renderer: Generative Object Rendering with 3D-aware Controllable Video Diffusion Models
por: Gu, Zekai, et al.
Publicado: (2026)
por: Gu, Zekai, et al.
Publicado: (2026)
Ranking-aware Continual Learning for LiDAR Place Recognition
por: Wang, Xufei, et al.
Publicado: (2025)
por: Wang, Xufei, et al.
Publicado: (2025)
CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
por: Song, Qi, et al.
Publicado: (2025)
por: Song, Qi, et al.
Publicado: (2025)
StimuVAR: Spatiotemporal Stimuli-aware Video Affective Reasoning with Multimodal Large Language Models
por: Guo, Yuxiang, et al.
Publicado: (2024)
por: Guo, Yuxiang, et al.
Publicado: (2024)
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
por: Liu, Chunxu, et al.
Publicado: (2025)
por: Liu, Chunxu, et al.
Publicado: (2025)
Unlocking Reasoning Potential in Large Langauge Models by Scaling Code-form Planning
por: Wen, Jiaxin, et al.
Publicado: (2024)
por: Wen, Jiaxin, et al.
Publicado: (2024)
A conjecture of Nadji, Ahmia and Ram\'ırez on congruences for biregular overpartitions
por: Tang, Dazhao
Publicado: (2025)
por: Tang, Dazhao
Publicado: (2025)
Simulation of lethal control and fertility control in a demographic model for Brandt's vole Microtus brandti. / Dazhao Shi
por: Shi, Dazhao
Publicado: (1995)
por: Shi, Dazhao
Publicado: (1995)
Reasoning-Enhanced Object-Centric Learning for Videos
por: Li, Jian, et al.
Publicado: (2024)
por: Li, Jian, et al.
Publicado: (2024)
Controllable Video Object Insertion via Multiview Priors
por: Qi, Xia, et al.
Publicado: (2026)
por: Qi, Xia, et al.
Publicado: (2026)
End-To-End Underwater Video Enhancement: Dataset and Model
por: Du, Dazhao, et al.
Publicado: (2024)
por: Du, Dazhao, et al.
Publicado: (2024)
Motion-aware Memory Network for Fast Video Salient Object Detection
por: Zhao, Xing, et al.
Publicado: (2022)
por: Zhao, Xing, et al.
Publicado: (2022)
Interaction-Consistent Object Removal via MLLM-Based Reasoning
por: Huang, Ching-Kai, et al.
Publicado: (2026)
por: Huang, Ching-Kai, et al.
Publicado: (2026)
Elysium: Exploring Object-level Perception in Videos via MLLM
por: Wang, Han, et al.
Publicado: (2024)
por: Wang, Han, et al.
Publicado: (2024)
Multi-MLLM Knowledge Distillation for Out-of-Context News Detection
por: Gu, Yimeng, et al.
Publicado: (2025)
por: Gu, Yimeng, et al.
Publicado: (2025)
SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
por: Chen, Xinyu, et al.
Publicado: (2026)
por: Chen, Xinyu, et al.
Publicado: (2026)
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
por: Guo, Meng-Hao, et al.
Publicado: (2025)
por: Guo, Meng-Hao, et al.
Publicado: (2025)
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
por: Zhao, Pengcheng, et al.
Publicado: (2024)
por: Zhao, Pengcheng, et al.
Publicado: (2024)
Design-MLLM: A Reinforcement Alignment Framework for Verifiable and Aesthetic Interior Design
por: Yang, Yuxuan, et al.
Publicado: (2026)
por: Yang, Yuxuan, et al.
Publicado: (2026)
Ejemplares similares
-
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
por: Gu, Bohai, et al.
Publicado: (2026) -
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
por: Du, Dazhao, et al.
Publicado: (2026) -
Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
por: Zhao, Lingxiao, et al.
Publicado: (2025) -
ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation
por: Yang, Songlin, et al.
Publicado: (2026) -
Coherent Video Inpainting Using Optical Flow-Guided Efficient Diffusion
por: Gu, Bohai, et al.
Publicado: (2024)