IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yunong, Eyzaguirre, Cristobal, Li, Manling, Khanna, Shubh, Niebles, Juan Carlos, Ravi, Vineeth, Mishra, Saumitra, Liu, Weiyu, Wu, Jiajun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Linear Scaling Video VLMs for Long Video Understanding
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via Vision-Language Models
von: Tie, Chenrui, et al.
Veröffentlicht: (2025)
von: Tie, Chenrui, et al.
Veröffentlicht: (2025)
Learning Planning Abstractions from Language
von: Liu, Weiyu, et al.
Veröffentlicht: (2024)
von: Liu, Weiyu, et al.
Veröffentlicht: (2024)
Hierarchical Language Models for Semantic Navigation and Manipulation in an Aerial-Ground Robotic System
von: Liu, Haokun, et al.
Veröffentlicht: (2025)
von: Liu, Haokun, et al.
Veröffentlicht: (2025)
Neural Elevation Models for Terrain Mapping and Path Planning
von: Dai, Adam, et al.
Veröffentlicht: (2024)
von: Dai, Adam, et al.
Veröffentlicht: (2024)
BrickCraft: Visuomotor Skill Composition with Situated Manual Guidance for Long-Horizon Interlocking Brick Assembly
von: Yu, Jichuan, et al.
Veröffentlicht: (2026)
von: Yu, Jichuan, et al.
Veröffentlicht: (2026)
Tactile Neural De-rendering
von: Eyzaguirre, Jose A., et al.
Veröffentlicht: (2024)
von: Eyzaguirre, Jose A., et al.
Veröffentlicht: (2024)
Streaming Detection of Queried Event Start
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2024)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2024)
Multimodal Sensing and Machine Learning to Compare Printed and Verbal Assembly Instructions Delivered by a Social Robot
von: Mishra, Ruchik, et al.
Veröffentlicht: (2025)
von: Mishra, Ruchik, et al.
Veröffentlicht: (2025)
Manual2Skill: Learning to Read Manuals and Acquire Robotic Skills for Furniture Assembly Using Vision-Language Models
von: Tie, Chenrui, et al.
Veröffentlicht: (2025)
von: Tie, Chenrui, et al.
Veröffentlicht: (2025)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
von: Ye, Jinhui, et al.
Veröffentlicht: (2025)
von: Ye, Jinhui, et al.
Veröffentlicht: (2025)
Learning Compositional Behaviors from Demonstration and Language
von: Liu, Weiyu, et al.
Veröffentlicht: (2025)
von: Liu, Weiyu, et al.
Veröffentlicht: (2025)
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
Understanding Complexity in VideoQA via Visual Program Generation
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025)
MetaWorld: Skill Transfer and Composition in a Hierarchical World Model for Grounding High-Level Instructions
von: Shen, Yutong, et al.
Veröffentlicht: (2026)
von: Shen, Yutong, et al.
Veröffentlicht: (2026)
Neural Radiance Maps for Extraterrestrial Navigation and Path Planning
von: Dai, Adam, et al.
Veröffentlicht: (2026)
von: Dai, Adam, et al.
Veröffentlicht: (2026)
Ground Penetrating Radar-Assisted Multimodal Robot Odometry Using Subsurface Feature Matrix
von: Li, Haifeng, et al.
Veröffentlicht: (2025)
von: Li, Haifeng, et al.
Veröffentlicht: (2025)
3D UAV Trajectory Estimation and Classification from Internet Videos via Language Model
von: Lei, Haoxiang, et al.
Veröffentlicht: (2026)
von: Lei, Haoxiang, et al.
Veröffentlicht: (2026)
AssemblyComplete: 3D Combinatorial Construction with Deep Reinforcement Learning
von: Chen, Alan, et al.
Veröffentlicht: (2024)
von: Chen, Alan, et al.
Veröffentlicht: (2024)
Using large language models for embodied planning introduces systematic safety risks
von: Zhang, Tao, et al.
Veröffentlicht: (2026)
von: Zhang, Tao, et al.
Veröffentlicht: (2026)
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
von: Lu, Haoran, et al.
Veröffentlicht: (2026)
von: Lu, Haoran, et al.
Veröffentlicht: (2026)
Composable Part-Based Manipulation
von: Liu, Weiyu, et al.
Veröffentlicht: (2024)
von: Liu, Weiyu, et al.
Veröffentlicht: (2024)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
von: Gu, Chenyang, et al.
Veröffentlicht: (2025)
von: Gu, Chenyang, et al.
Veröffentlicht: (2025)
Enhancing Autonomous Manipulator Control with Human-in-loop for Uncertain Assembly Environments
von: Mishra, Ashutosh, et al.
Veröffentlicht: (2025)
von: Mishra, Ashutosh, et al.
Veröffentlicht: (2025)
Eye-in-Finger: Smart Fingers for Delicate Assembly and Disassembly of LEGO
von: Tang, Zhenran, et al.
Veröffentlicht: (2025)
von: Tang, Zhenran, et al.
Veröffentlicht: (2025)
StableLego: Stability Analysis of Block Stacking Assembly
von: Liu, Ruixuan, et al.
Veröffentlicht: (2024)
von: Liu, Ruixuan, et al.
Veröffentlicht: (2024)
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
A Model-Agnostic Approach for Semantically Driven Disambiguation in Human-Robot Interaction
von: Dogan, Fethiye Irmak, et al.
Veröffentlicht: (2024)
von: Dogan, Fethiye Irmak, et al.
Veröffentlicht: (2024)
Latent Space Planning for Multi-Object Manipulation with Environment-Aware Relational Classifiers
von: Huang, Yixuan, et al.
Veröffentlicht: (2023)
von: Huang, Yixuan, et al.
Veröffentlicht: (2023)
Point What You Mean: Visually Grounded Instruction Policy
von: Yu, Hang, et al.
Veröffentlicht: (2025)
von: Yu, Hang, et al.
Veröffentlicht: (2025)
Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos
von: Ye, Weirui, et al.
Veröffentlicht: (2025)
von: Ye, Weirui, et al.
Veröffentlicht: (2025)
Video-to-BT: Generating Reactive Behavior Trees from Human Demonstration Videos for Robotic Assembly
von: Zhao, Xiwei, et al.
Veröffentlicht: (2025)
von: Zhao, Xiwei, et al.
Veröffentlicht: (2025)
6D Object Pose Tracking in Internet Videos for Robotic Manipulation
von: Ponimatkin, Georgy, et al.
Veröffentlicht: (2025)
von: Ponimatkin, Georgy, et al.
Veröffentlicht: (2025)
Robots that Learn to Safely Influence via Prediction-Informed Reach-Avoid Dynamic Games
von: Pandya, Ravi, et al.
Veröffentlicht: (2024)
von: Pandya, Ravi, et al.
Veröffentlicht: (2024)
Multimodal Safe Control for Human-Robot Interaction
von: Pandya, Ravi, et al.
Veröffentlicht: (2023)
von: Pandya, Ravi, et al.
Veröffentlicht: (2023)
APEX-MR: Multi-Robot Asynchronous Planning and Execution for Cooperative Assembly
von: Huang, Philip, et al.
Veröffentlicht: (2025)
von: Huang, Philip, et al.
Veröffentlicht: (2025)
On Accurate and Robust Estimation of 3D and 2D Circular Center: Method and Application to Camera-Lidar Calibration
von: Jiang, Jiajun, et al.
Veröffentlicht: (2025)
von: Jiang, Jiajun, et al.
Veröffentlicht: (2025)
Taming generative video models for zero-shot optical flow extraction
von: Kim, Seungwoo, et al.
Veröffentlicht: (2025)
von: Kim, Seungwoo, et al.
Veröffentlicht: (2025)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
von: Li, Manling, et al.
Veröffentlicht: (2024)
von: Li, Manling, et al.
Veröffentlicht: (2024)
Robot voice a voice controlled robot using arduino
von: Teeda, Vineeth, et al.
Veröffentlicht: (2024)
von: Teeda, Vineeth, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Linear Scaling Video VLMs for Long Video Understanding
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026) -
Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via Vision-Language Models
von: Tie, Chenrui, et al.
Veröffentlicht: (2025) -
Learning Planning Abstractions from Language
von: Liu, Weiyu, et al.
Veröffentlicht: (2024) -
Hierarchical Language Models for Semantic Navigation and Manipulation in an Aerial-Ground Robotic System
von: Liu, Haokun, et al.
Veröffentlicht: (2025) -
Neural Elevation Models for Terrain Mapping and Path Planning
von: Dai, Adam, et al.
Veröffentlicht: (2024)