H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ci, Hai, Liu, Xiaokang, Yang, Pei, Song, Yiren, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
von: Yang, Pei, et al.
Veröffentlicht: (2025)
von: Yang, Pei, et al.
Veröffentlicht: (2025)
RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
von: Mei, Haiyang, et al.
Veröffentlicht: (2025)
von: Mei, Haiyang, et al.
Veröffentlicht: (2025)
UENR-600K: A Large-Scale Physically Grounded Dataset for Nighttime Video Deraining
von: Yang, Pei, et al.
Veröffentlicht: (2026)
von: Yang, Pei, et al.
Veröffentlicht: (2026)
World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy
von: Liu, Xiaokang, et al.
Veröffentlicht: (2026)
von: Liu, Xiaokang, et al.
Veröffentlicht: (2026)
OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
Steganalysis on Digital Watermarking: Is Your Defense Truly Impervious?
von: Yang, Pei, et al.
Veröffentlicht: (2024)
von: Yang, Pei, et al.
Veröffentlicht: (2024)
IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation
von: Song, Yiren, et al.
Veröffentlicht: (2024)
von: Song, Yiren, et al.
Veröffentlicht: (2024)
RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-Key Identification
von: Ci, Hai, et al.
Veröffentlicht: (2024)
von: Ci, Hai, et al.
Veröffentlicht: (2024)
Mitty: Diffusion-based Human-to-Robot Video Generation
von: Song, Yiren, et al.
Veröffentlicht: (2025)
von: Song, Yiren, et al.
Veröffentlicht: (2025)
Impossible Videos
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
Human2Robot: Learning Robot Actions from Paired Human-Robot Videos
von: Xie, Sicheng, et al.
Veröffentlicht: (2025)
von: Xie, Sicheng, et al.
Veröffentlicht: (2025)
H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
von: Li, Guangrun, et al.
Veröffentlicht: (2025)
von: Li, Guangrun, et al.
Veröffentlicht: (2025)
macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
von: Yang, Pei, et al.
Veröffentlicht: (2025)
von: Yang, Pei, et al.
Veröffentlicht: (2025)
WMAdapter: Adding WaterMark Control to Latent Diffusion Models
von: Ci, Hai, et al.
Veröffentlicht: (2024)
von: Ci, Hai, et al.
Veröffentlicht: (2024)
Anti-Reference: Universal and Immediate Defense Against Reference-Based Generation
von: Song, Yiren, et al.
Veröffentlicht: (2024)
von: Song, Yiren, et al.
Veröffentlicht: (2024)
Human-to-Robot Interaction: Learning from Video Demonstration for Robot Imitation
von: Canh, Thanh Nguyen, et al.
Veröffentlicht: (2026)
von: Canh, Thanh Nguyen, et al.
Veröffentlicht: (2026)
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
von: Song, Yiren, et al.
Veröffentlicht: (2025)
von: Song, Yiren, et al.
Veröffentlicht: (2025)
DiffSim: Taming Diffusion Models for Evaluating Visual Similarity
von: Song, Yiren, et al.
Veröffentlicht: (2024)
von: Song, Yiren, et al.
Veröffentlicht: (2024)
StreamingEffect: Real-Time Human-Centric Video Effect Generation
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
In-Context Defense in Computer Agents: An Empirical Study
von: Yang, Pei, et al.
Veröffentlicht: (2025)
von: Yang, Pei, et al.
Veröffentlicht: (2025)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
von: Xu, Runsen, et al.
Veröffentlicht: (2024)
von: Xu, Runsen, et al.
Veröffentlicht: (2024)
OmniUMI: Towards Physically Grounded Robot Learning via Human-Aligned Multimodal Interaction
von: Luo, Shaqi, et al.
Veröffentlicht: (2026)
von: Luo, Shaqi, et al.
Veröffentlicht: (2026)
Bidirectional Human-Robot Communication for Physical Human-Robot Interaction
von: Wang, Junxiang, et al.
Veröffentlicht: (2026)
von: Wang, Junxiang, et al.
Veröffentlicht: (2026)
An Approach to Combining Video and Speech with Large Language Models in Human-Robot Interaction
von: Shen, Guanting, et al.
Veröffentlicht: (2026)
von: Shen, Guanting, et al.
Veröffentlicht: (2026)
AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models
von: Huynh, Cuong, et al.
Veröffentlicht: (2026)
von: Huynh, Cuong, et al.
Veröffentlicht: (2026)
In-situ Value-aligned Human-Robot Interactions with Physical Constraints
von: Li, Hongtao, et al.
Veröffentlicht: (2025)
von: Li, Hongtao, et al.
Veröffentlicht: (2025)
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
von: Song, Zijian, et al.
Veröffentlicht: (2026)
von: Song, Zijian, et al.
Veröffentlicht: (2026)
Phantom: Training Robots Without Robots Using Only Human Videos
von: Lepert, Marion, et al.
Veröffentlicht: (2025)
von: Lepert, Marion, et al.
Veröffentlicht: (2025)
VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
von: Sun, Xuefei, et al.
Veröffentlicht: (2026)
von: Sun, Xuefei, et al.
Veröffentlicht: (2026)
EasyMimic: A Low-Cost Framework for Robot Imitation Learning from Human Videos
von: Zhang, Tao, et al.
Veröffentlicht: (2026)
von: Zhang, Tao, et al.
Veröffentlicht: (2026)
From Generated Human Videos to Physically Plausible Robot Trajectories
von: Ni, James, et al.
Veröffentlicht: (2025)
von: Ni, James, et al.
Veröffentlicht: (2025)
Video-to-BT: Generating Reactive Behavior Trees from Human Demonstration Videos for Robotic Assembly
von: Zhao, Xiwei, et al.
Veröffentlicht: (2025)
von: Zhao, Xiwei, et al.
Veröffentlicht: (2025)
Morphology-Consistent Humanoid Interaction through Robot-Centric Video Synthesis
von: Xu, Weisheng, et al.
Veröffentlicht: (2026)
von: Xu, Weisheng, et al.
Veröffentlicht: (2026)
TC-IDM: Grounding Video Generation for Executable Zero-shot Robot Motion
von: Mi, Weishi, et al.
Veröffentlicht: (2026)
von: Mi, Weishi, et al.
Veröffentlicht: (2026)
Video-Enhanced Offline Reinforcement Learning: A Model-Based Approach
von: Pan, Minting, et al.
Veröffentlicht: (2025)
von: Pan, Minting, et al.
Veröffentlicht: (2025)
Task Adaptation in Industrial Human-Robot Interaction: Leveraging Riemannian Motion Policies
von: Allenspach, Mike, et al.
Veröffentlicht: (2024)
von: Allenspach, Mike, et al.
Veröffentlicht: (2024)
Unidirectional Human-Robot-Human Physical Interaction for Gait Training
von: Amato, Lorenzo, et al.
Veröffentlicht: (2024)
von: Amato, Lorenzo, et al.
Veröffentlicht: (2024)
Architectural HRI: Towards a Robotic Paradigm Shift in Human-Building Interaction
von: Nguyen, Alex Binh Vinh Duc
Veröffentlicht: (2026)
von: Nguyen, Alex Binh Vinh Duc
Veröffentlicht: (2026)
Using 3-D LiDAR Data for Safe Physical Human-Robot Interaction
von: Arora, Sarthak, et al.
Veröffentlicht: (2024)
von: Arora, Sarthak, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
von: Yang, Pei, et al.
Veröffentlicht: (2025) -
RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
von: Mei, Haiyang, et al.
Veröffentlicht: (2025) -
UENR-600K: A Large-Scale Physically Grounded Dataset for Nighttime Video Deraining
von: Yang, Pei, et al.
Veröffentlicht: (2026) -
World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy
von: Liu, Xiaokang, et al.
Veröffentlicht: (2026) -
OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
von: Song, Yiren, et al.
Veröffentlicht: (2026)