TC-IDM: Grounding Video Generation for Executable Zero-shot Robot Motion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mi, Weishi, Bao, Yong, Chi, Xiaowei, Ju, Xiaozhu, Qin, Zhiyuan, Ge, Kuangzhi, Tang, Kai, Jia, Peidong, Zhang, Shanghang, Tang, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test
von: Fan, Chun-Kai, et al.
Veröffentlicht: (2026)
von: Fan, Chun-Kai, et al.
Veröffentlicht: (2026)
WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation
von: Qian, Zezhong, et al.
Veröffentlicht: (2025)
von: Qian, Zezhong, et al.
Veröffentlicht: (2025)
WoW: Towards a World omniscient World model Through Embodied Interaction
von: Chi, Xiaowei, et al.
Veröffentlicht: (2025)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2025)
Can World Models Benefit VLMs for World Dynamics?
von: Zhang, Kevin, et al.
Veröffentlicht: (2025)
von: Zhang, Kevin, et al.
Veröffentlicht: (2025)
LABSHIELD: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories
von: Sun, Qianpu, et al.
Veröffentlicht: (2026)
von: Sun, Qianpu, et al.
Veröffentlicht: (2026)
MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis
von: Chi, Xiaowei, et al.
Veröffentlicht: (2025)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2025)
Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
von: Jia, Yueru, et al.
Veröffentlicht: (2025)
von: Jia, Yueru, et al.
Veröffentlicht: (2025)
SCBench: A Sports Commentary Benchmark for Video LLMs
von: Ge, Kuangzhi, et al.
Veröffentlicht: (2024)
von: Ge, Kuangzhi, et al.
Veröffentlicht: (2024)
SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents
von: Tian, Wanxin, et al.
Veröffentlicht: (2025)
von: Tian, Wanxin, et al.
Veröffentlicht: (2025)
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
Efficient Training of Generalizable Visuomotor Policies via Control-Aware Augmentation
von: Zhao, Yinuo, et al.
Veröffentlicht: (2024)
von: Zhao, Yinuo, et al.
Veröffentlicht: (2024)
Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2023)
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2023)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
von: Chen, Yi, et al.
Veröffentlicht: (2024)
von: Chen, Yi, et al.
Veröffentlicht: (2024)
Swampland bound on quintessential inflation in IDM
von: Saoud, S., et al.
Veröffentlicht: (2025)
von: Saoud, S., et al.
Veröffentlicht: (2025)
OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection
von: Liu, Yu, et al.
Veröffentlicht: (2025)
von: Liu, Yu, et al.
Veröffentlicht: (2025)
Zeus: Zero-shot LLM Instruction for Union Segmentation in Multimodal Medical Imaging
von: Dai, Siyuan, et al.
Veröffentlicht: (2025)
von: Dai, Siyuan, et al.
Veröffentlicht: (2025)
Combining Knowledge Graph and LLMs for Enhanced Zero-shot Visual Question Answering
von: Tao, Qian, et al.
Veröffentlicht: (2025)
von: Tao, Qian, et al.
Veröffentlicht: (2025)
TEST-V: TEst-time Support-set Tuning for Zero-shot Video Classification
von: Yan, Rui, et al.
Veröffentlicht: (2025)
von: Yan, Rui, et al.
Veröffentlicht: (2025)
ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
von: Li, Ying, et al.
Veröffentlicht: (2025)
von: Li, Ying, et al.
Veröffentlicht: (2025)
Extended IDM theory with low scale seesaw mechanisms
von: Huong, D. T., et al.
Veröffentlicht: (2025)
von: Huong, D. T., et al.
Veröffentlicht: (2025)
Collaborative Control Method of Transit Signal Priority Based on Cooperative Game and Reinforcement Learning
von: Qin, Hao, et al.
Veröffentlicht: (2024)
von: Qin, Hao, et al.
Veröffentlicht: (2024)
InstantID: Zero-shot Identity-Preserving Generation in Seconds
von: Wang, Qixun, et al.
Veröffentlicht: (2024)
von: Wang, Qixun, et al.
Veröffentlicht: (2024)
Zero-shot Prompt-based Video Encoder for Surgical Gesture Recognition
von: Rao, Mingxing, et al.
Veröffentlicht: (2024)
von: Rao, Mingxing, et al.
Veröffentlicht: (2024)
Diagnosing Shortcut-Induced Rigidity in Continual Learning: The Einstellung Rigidity Index (ERI)
von: Gu, Kai, et al.
Veröffentlicht: (2025)
von: Gu, Kai, et al.
Veröffentlicht: (2025)
PoseDiff: A Unified Diffusion Model Bridging Robot Pose Estimation and Video-to-Action Control
von: Zhang, Haozhuo, et al.
Veröffentlicht: (2025)
von: Zhang, Haozhuo, et al.
Veröffentlicht: (2025)
Large Motion Video Autoencoding with Cross-modal Video VAE
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
Segment Any Motion in Videos
von: Huang, Nan, et al.
Veröffentlicht: (2025)
von: Huang, Nan, et al.
Veröffentlicht: (2025)
Tex4D: Zero-shot 4D Scene Texturing with Video Diffusion Models
von: Bao, Jingzhi, et al.
Veröffentlicht: (2024)
von: Bao, Jingzhi, et al.
Veröffentlicht: (2024)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model
von: Wu, Pengying, et al.
Veröffentlicht: (2024)
von: Wu, Pengying, et al.
Veröffentlicht: (2024)
EVA: An Embodied World Model for Future Video Anticipation
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
Depth-aware Test-Time Training for Zero-shot Video Object Segmentation
von: Liu, Weihuang, et al.
Veröffentlicht: (2024)
von: Liu, Weihuang, et al.
Veröffentlicht: (2024)
Few-shot Class-Incremental Learning via Generative Co-Memory Regularization
von: Bao, Kexin, et al.
Veröffentlicht: (2026)
von: Bao, Kexin, et al.
Veröffentlicht: (2026)
Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data
von: Fan, Ke, et al.
Veröffentlicht: (2025)
von: Fan, Ke, et al.
Veröffentlicht: (2025)
Improving Zero-shot LLM Re-Ranker with Risk Minimization
von: Yuan, Xiaowei, et al.
Veröffentlicht: (2024)
von: Yuan, Xiaowei, et al.
Veröffentlicht: (2024)
System Synchronization Based on Complex Frequency
von: Wei, Yusen, et al.
Veröffentlicht: (2025)
von: Wei, Yusen, et al.
Veröffentlicht: (2025)
Fine-gained Zero-shot Video Sampling
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis
von: Wang, Yuji, et al.
Veröffentlicht: (2025)
von: Wang, Yuji, et al.
Veröffentlicht: (2025)
FastInit: Fast Noise Initialization for Temporally Consistent Video Generation
von: Bai, Chengyu, et al.
Veröffentlicht: (2025)
von: Bai, Chengyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test
von: Fan, Chun-Kai, et al.
Veröffentlicht: (2026) -
WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation
von: Qian, Zezhong, et al.
Veröffentlicht: (2025) -
WoW: Towards a World omniscient World model Through Embodied Interaction
von: Chi, Xiaowei, et al.
Veröffentlicht: (2025) -
Can World Models Benefit VLMs for World Dynamics?
von: Zhang, Kevin, et al.
Veröffentlicht: (2025) -
LABSHIELD: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories
von: Sun, Qianpu, et al.
Veröffentlicht: (2026)