Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zahid, Azizul, Fan, Jie, Wang, Farong, Dy, Ashton, Swaminathan, Sai, Liu, Fei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
arg-VU: Affordance Reasoning with Physics-Aware 3D Geometry for Visual Understanding in Robotic Surgery
von: Xiao, Nan, et al.
Veröffentlicht: (2026)
von: Xiao, Nan, et al.
Veröffentlicht: (2026)
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
von: Jiang, Yuming, et al.
Veröffentlicht: (2025)
von: Jiang, Yuming, et al.
Veröffentlicht: (2025)
GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and Imitation
von: Wang, Zifan, et al.
Veröffentlicht: (2024)
von: Wang, Zifan, et al.
Veröffentlicht: (2024)
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
von: Xiao, Zhanqi, et al.
Veröffentlicht: (2026)
von: Xiao, Zhanqi, et al.
Veröffentlicht: (2026)
MEM: Multi-Modal Elevation Mapping for Robotics and Learning
von: Erni, Gian, et al.
Veröffentlicht: (2023)
von: Erni, Gian, et al.
Veröffentlicht: (2023)
Multi-Modal Graph Convolutional Network with Sinusoidal Encoding for Robust Human Action Segmentation
von: Xing, Hao, et al.
Veröffentlicht: (2025)
von: Xing, Hao, et al.
Veröffentlicht: (2025)
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
von: Wang, Kuanning, et al.
Veröffentlicht: (2026)
von: Wang, Kuanning, et al.
Veröffentlicht: (2026)
ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
von: Zeng, Qiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Qiyuan, et al.
Veröffentlicht: (2025)
Surface-Constrained Offline Warping with Contact-Aware Online Pose Projection for Safe Robotic Trajectory Execution
von: Wang, Farong, et al.
Veröffentlicht: (2026)
von: Wang, Farong, et al.
Veröffentlicht: (2026)
A Distributed Multi-Modal Sensing Approach for Human Activity Recognition in Real-Time Human-Robot Collaboration
von: Belcamino, Valerio, et al.
Veröffentlicht: (2026)
von: Belcamino, Valerio, et al.
Veröffentlicht: (2026)
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
von: Yin, Zhenhan, et al.
Veröffentlicht: (2025)
von: Yin, Zhenhan, et al.
Veröffentlicht: (2025)
Continual Learning for Autonomous Robots: A Prototype-based Approach
von: Hajizada, Elvin, et al.
Veröffentlicht: (2024)
von: Hajizada, Elvin, et al.
Veröffentlicht: (2024)
Rodrigues Network for Learning Robot Actions
von: Zhang, Jialiang, et al.
Veröffentlicht: (2025)
von: Zhang, Jialiang, et al.
Veröffentlicht: (2025)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement
von: Wang, Kohou, et al.
Veröffentlicht: (2025)
von: Wang, Kohou, et al.
Veröffentlicht: (2025)
You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations
von: Zhou, Huayi, et al.
Veröffentlicht: (2025)
von: Zhou, Huayi, et al.
Veröffentlicht: (2025)
Deep Learning-Based Multi-Modal Fusion for Robust Robot Perception and Navigation
von: Lai, Delun, et al.
Veröffentlicht: (2025)
von: Lai, Delun, et al.
Veröffentlicht: (2025)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026)
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface
von: Zhao, Yujie, et al.
Veröffentlicht: (2025)
von: Zhao, Yujie, et al.
Veröffentlicht: (2025)
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
von: Lei, Zixing, et al.
Veröffentlicht: (2026)
von: Lei, Zixing, et al.
Veröffentlicht: (2026)
DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
von: Bai, Yang, et al.
Veröffentlicht: (2025)
von: Bai, Yang, et al.
Veröffentlicht: (2025)
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
von: Huang, Zheng, et al.
Veröffentlicht: (2025)
von: Huang, Zheng, et al.
Veröffentlicht: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
Towards Open-World Human Action Segmentation Using Graph Convolutional Networks
von: Xing, Hao, et al.
Veröffentlicht: (2025)
von: Xing, Hao, et al.
Veröffentlicht: (2025)
Self-supervised 6-DoF Robot Grasping by Demonstration via Augmented Reality Teleoperation System
von: Dengxiong, Xiwen, et al.
Veröffentlicht: (2024)
von: Dengxiong, Xiwen, et al.
Veröffentlicht: (2024)
WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
von: Guan, Runwei, et al.
Veröffentlicht: (2026)
von: Guan, Runwei, et al.
Veröffentlicht: (2026)
Towards Dense and Accurate Radar Perception Via Efficient Cross-Modal Diffusion Model
von: Zhang, Ruibin, et al.
Veröffentlicht: (2024)
von: Zhang, Ruibin, et al.
Veröffentlicht: (2024)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
von: Guo, Jianing, et al.
Veröffentlicht: (2025)
von: Guo, Jianing, et al.
Veröffentlicht: (2025)
How Physics and Background Attributes Impact Video Transformers in Robotic Manipulation: A Case Study on Planar Pushing
von: Jin, Shutong, et al.
Veröffentlicht: (2023)
von: Jin, Shutong, et al.
Veröffentlicht: (2023)
Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations
von: Huang, Wei-Jin, et al.
Veröffentlicht: (2026)
von: Huang, Wei-Jin, et al.
Veröffentlicht: (2026)
Generate, Transfer, Adapt: Learning Functional Dexterous Grasping from a Single Human Demonstration
von: He, Xingyi, et al.
Veröffentlicht: (2026)
von: He, Xingyi, et al.
Veröffentlicht: (2026)
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
von: Chen, Hanzhi, et al.
Veröffentlicht: (2025)
von: Chen, Hanzhi, et al.
Veröffentlicht: (2025)
SCAR: Self-Supervised Continuous Action Representation Learning
von: Liu, Hongjia, et al.
Veröffentlicht: (2026)
von: Liu, Hongjia, et al.
Veröffentlicht: (2026)
Multi-Camera Hand-Eye Calibration for Human-Robot Collaboration in Industrial Robotic Workcells
von: Allegro, Davide, et al.
Veröffentlicht: (2024)
von: Allegro, Davide, et al.
Veröffentlicht: (2024)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
arg-VU: Affordance Reasoning with Physics-Aware 3D Geometry for Visual Understanding in Robotic Surgery
von: Xiao, Nan, et al.
Veröffentlicht: (2026) -
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
von: Jiang, Yuming, et al.
Veröffentlicht: (2025) -
GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and Imitation
von: Wang, Zifan, et al.
Veröffentlicht: (2024) -
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
von: Xiao, Zhanqi, et al.
Veröffentlicht: (2026) -
MEM: Multi-Modal Elevation Mapping for Robotics and Learning
von: Erni, Gian, et al.
Veröffentlicht: (2023)