Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Zahid, Azizul, Fan, Jie, Wang, Farong, Dy, Ashton, Swaminathan, Sai, Liu, Fei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
arg-VU: Affordance Reasoning with Physics-Aware 3D Geometry for Visual Understanding in Robotic Surgery
di: Xiao, Nan, et al.
Pubblicazione: (2026)
di: Xiao, Nan, et al.
Pubblicazione: (2026)
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
di: Jiang, Yuming, et al.
Pubblicazione: (2025)
di: Jiang, Yuming, et al.
Pubblicazione: (2025)
GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and Imitation
di: Wang, Zifan, et al.
Pubblicazione: (2024)
di: Wang, Zifan, et al.
Pubblicazione: (2024)
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
di: Xiao, Zhanqi, et al.
Pubblicazione: (2026)
di: Xiao, Zhanqi, et al.
Pubblicazione: (2026)
MEM: Multi-Modal Elevation Mapping for Robotics and Learning
di: Erni, Gian, et al.
Pubblicazione: (2023)
di: Erni, Gian, et al.
Pubblicazione: (2023)
Multi-Modal Graph Convolutional Network with Sinusoidal Encoding for Robust Human Action Segmentation
di: Xing, Hao, et al.
Pubblicazione: (2025)
di: Xing, Hao, et al.
Pubblicazione: (2025)
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
di: Wang, Kuanning, et al.
Pubblicazione: (2026)
di: Wang, Kuanning, et al.
Pubblicazione: (2026)
ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
di: Zeng, Qiyuan, et al.
Pubblicazione: (2025)
di: Zeng, Qiyuan, et al.
Pubblicazione: (2025)
Surface-Constrained Offline Warping with Contact-Aware Online Pose Projection for Safe Robotic Trajectory Execution
di: Wang, Farong, et al.
Pubblicazione: (2026)
di: Wang, Farong, et al.
Pubblicazione: (2026)
A Distributed Multi-Modal Sensing Approach for Human Activity Recognition in Real-Time Human-Robot Collaboration
di: Belcamino, Valerio, et al.
Pubblicazione: (2026)
di: Belcamino, Valerio, et al.
Pubblicazione: (2026)
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
di: Li, Wei, et al.
Pubblicazione: (2025)
di: Li, Wei, et al.
Pubblicazione: (2025)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
di: Yin, Zhenhan, et al.
Pubblicazione: (2025)
di: Yin, Zhenhan, et al.
Pubblicazione: (2025)
Continual Learning for Autonomous Robots: A Prototype-based Approach
di: Hajizada, Elvin, et al.
Pubblicazione: (2024)
di: Hajizada, Elvin, et al.
Pubblicazione: (2024)
Rodrigues Network for Learning Robot Actions
di: Zhang, Jialiang, et al.
Pubblicazione: (2025)
di: Zhang, Jialiang, et al.
Pubblicazione: (2025)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
di: Wang, Ziyi, et al.
Pubblicazione: (2025)
di: Wang, Ziyi, et al.
Pubblicazione: (2025)
iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement
di: Wang, Kohou, et al.
Pubblicazione: (2025)
di: Wang, Kohou, et al.
Pubblicazione: (2025)
You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations
di: Zhou, Huayi, et al.
Pubblicazione: (2025)
di: Zhou, Huayi, et al.
Pubblicazione: (2025)
Deep Learning-Based Multi-Modal Fusion for Robust Robot Perception and Navigation
di: Lai, Delun, et al.
Pubblicazione: (2025)
di: Lai, Delun, et al.
Pubblicazione: (2025)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
di: Wen, Junjie, et al.
Pubblicazione: (2024)
di: Wen, Junjie, et al.
Pubblicazione: (2024)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
di: Liu, Yicheng, et al.
Pubblicazione: (2025)
di: Liu, Yicheng, et al.
Pubblicazione: (2025)
Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface
di: Zhao, Yujie, et al.
Pubblicazione: (2025)
di: Zhao, Yujie, et al.
Pubblicazione: (2025)
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
di: Lei, Zixing, et al.
Pubblicazione: (2026)
di: Lei, Zixing, et al.
Pubblicazione: (2026)
DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
di: Bai, Yang, et al.
Pubblicazione: (2025)
di: Bai, Yang, et al.
Pubblicazione: (2025)
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
di: Huang, Zheng, et al.
Pubblicazione: (2025)
di: Huang, Zheng, et al.
Pubblicazione: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
di: Shi, Hao, et al.
Pubblicazione: (2025)
di: Shi, Hao, et al.
Pubblicazione: (2025)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
Towards Open-World Human Action Segmentation Using Graph Convolutional Networks
di: Xing, Hao, et al.
Pubblicazione: (2025)
di: Xing, Hao, et al.
Pubblicazione: (2025)
Self-supervised 6-DoF Robot Grasping by Demonstration via Augmented Reality Teleoperation System
di: Dengxiong, Xiwen, et al.
Pubblicazione: (2024)
di: Dengxiong, Xiwen, et al.
Pubblicazione: (2024)
WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
di: Guan, Runwei, et al.
Pubblicazione: (2026)
di: Guan, Runwei, et al.
Pubblicazione: (2026)
Towards Dense and Accurate Radar Perception Via Efficient Cross-Modal Diffusion Model
di: Zhang, Ruibin, et al.
Pubblicazione: (2024)
di: Zhang, Ruibin, et al.
Pubblicazione: (2024)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
di: Shi, Jiaqi, et al.
Pubblicazione: (2026)
di: Shi, Jiaqi, et al.
Pubblicazione: (2026)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
di: Guo, Jianing, et al.
Pubblicazione: (2025)
di: Guo, Jianing, et al.
Pubblicazione: (2025)
How Physics and Background Attributes Impact Video Transformers in Robotic Manipulation: A Case Study on Planar Pushing
di: Jin, Shutong, et al.
Pubblicazione: (2023)
di: Jin, Shutong, et al.
Pubblicazione: (2023)
Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations
di: Huang, Wei-Jin, et al.
Pubblicazione: (2026)
di: Huang, Wei-Jin, et al.
Pubblicazione: (2026)
Generate, Transfer, Adapt: Learning Functional Dexterous Grasping from a Single Human Demonstration
di: He, Xingyi, et al.
Pubblicazione: (2026)
di: He, Xingyi, et al.
Pubblicazione: (2026)
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
di: Chen, Hanzhi, et al.
Pubblicazione: (2025)
di: Chen, Hanzhi, et al.
Pubblicazione: (2025)
SCAR: Self-Supervised Continuous Action Representation Learning
di: Liu, Hongjia, et al.
Pubblicazione: (2026)
di: Liu, Hongjia, et al.
Pubblicazione: (2026)
Multi-Camera Hand-Eye Calibration for Human-Robot Collaboration in Industrial Robotic Workcells
di: Allegro, Davide, et al.
Pubblicazione: (2024)
di: Allegro, Davide, et al.
Pubblicazione: (2024)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
di: Mitra, Chancharik, et al.
Pubblicazione: (2025)
di: Mitra, Chancharik, et al.
Pubblicazione: (2025)
Documenti analoghi
-
arg-VU: Affordance Reasoning with Physics-Aware 3D Geometry for Visual Understanding in Robotic Surgery
di: Xiao, Nan, et al.
Pubblicazione: (2026) -
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
di: Jiang, Yuming, et al.
Pubblicazione: (2025) -
GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and Imitation
di: Wang, Zifan, et al.
Pubblicazione: (2024) -
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
di: Xiao, Zhanqi, et al.
Pubblicazione: (2026) -
MEM: Multi-Modal Elevation Mapping for Robotics and Learning
di: Erni, Gian, et al.
Pubblicazione: (2023)