HumanNet: Scaling Human-centric Video Learning to One Million Hours
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Yufan, Zhou, Daquan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Video Generation Model for the Embodied World
by: Deng, Yufan, et al.
Published: (2026)
by: Deng, Yufan, et al.
Published: (2026)
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
by: Fu, Yiyang, et al.
Published: (2026)
by: Fu, Yiyang, et al.
Published: (2026)
Robot Learning from Human Videos: A Survey
by: Ma, Junyi, et al.
Published: (2026)
by: Ma, Junyi, et al.
Published: (2026)
You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations
by: Zhou, Huayi, et al.
Published: (2025)
by: Zhou, Huayi, et al.
Published: (2025)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
Object-centric 3D Motion Field for Robot Learning from Human Videos
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
CEDex: Cross-Embodiment Dexterous Grasp Generation at Scale from Human-like Contact Representations
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories
by: Mei, Yanghong, et al.
Published: (2025)
by: Mei, Yanghong, et al.
Published: (2025)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
AoE: Always-on Egocentric Human Video Collection for Embodied AI
by: Yang, Bowen, et al.
Published: (2026)
by: Yang, Bowen, et al.
Published: (2026)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
Kimodo: Scaling Controllable Human Motion Generation
by: Rempe, Davis, et al.
Published: (2026)
by: Rempe, Davis, et al.
Published: (2026)
UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
by: Kim, Hanjung, et al.
Published: (2025)
by: Kim, Hanjung, et al.
Published: (2025)
GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
by: Ma, Teli, et al.
Published: (2025)
by: Ma, Teli, et al.
Published: (2025)
Learning Priors of Human Motion With Vision Transformers
by: Falqueto, Placido, et al.
Published: (2025)
by: Falqueto, Placido, et al.
Published: (2025)
One-shot Video Imitation via Parameterized Symbolic Abstraction Graphs
by: Wang, Jianren, et al.
Published: (2024)
by: Wang, Jianren, et al.
Published: (2024)
Slot-Level Robotic Placement via Visual Imitation from Single Human Video
by: Shan, Dandan, et al.
Published: (2025)
by: Shan, Dandan, et al.
Published: (2025)
HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
by: Luo, Hao, et al.
Published: (2025)
by: Luo, Hao, et al.
Published: (2025)
HRDexDB: A Large-Scale Dataset of Dexterous Human and Robotic Hand Grasps
by: Lim, Jongbin, et al.
Published: (2026)
by: Lim, Jongbin, et al.
Published: (2026)
EgoMimic: Scaling Imitation Learning via Egocentric Video
by: Kareer, Simar, et al.
Published: (2024)
by: Kareer, Simar, et al.
Published: (2024)
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
by: Chen, Hanzhi, et al.
Published: (2025)
by: Chen, Hanzhi, et al.
Published: (2025)
Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning
by: Shibata, Yuto, et al.
Published: (2026)
by: Shibata, Yuto, et al.
Published: (2026)
Generate, Transfer, Adapt: Learning Functional Dexterous Grasping from a Single Human Demonstration
by: He, Xingyi, et al.
Published: (2026)
by: He, Xingyi, et al.
Published: (2026)
Bayesian-Optimized One-Step Diffusion Model with Knowledge Distillation for Real-Time 3D Human Motion Prediction
by: Tian, Sibo, et al.
Published: (2024)
by: Tian, Sibo, et al.
Published: (2024)
DexMan: Learning Bimanual Dexterous Manipulation from Human and Generated Videos
by: Hsieh, Jhen, et al.
Published: (2025)
by: Hsieh, Jhen, et al.
Published: (2025)
Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations
by: Huang, Wei-Jin, et al.
Published: (2026)
by: Huang, Wei-Jin, et al.
Published: (2026)
Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction
by: Chen, Hongyi, et al.
Published: (2026)
by: Chen, Hongyi, et al.
Published: (2026)
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
by: Liu, Xinhao, et al.
Published: (2024)
by: Liu, Xinhao, et al.
Published: (2024)
VizFlyt: Perception-centric Pedagogical Framework For Autonomous Aerial Robots
by: Srivastava, Kushagra, et al.
Published: (2025)
by: Srivastava, Kushagra, et al.
Published: (2025)
MetricNet: Recovering Metric Scale in Generative Navigation Policies
by: Nayak, Abhijeet, et al.
Published: (2025)
by: Nayak, Abhijeet, et al.
Published: (2025)
HHI-Assist: A Dataset and Benchmark of Human-Human Interaction in Physical Assistance Scenario
by: Saadatnejad, Saeed, et al.
Published: (2025)
by: Saadatnejad, Saeed, et al.
Published: (2025)
HOH: Markerless Multimodal Human-Object-Human Handover Dataset with Large Object Count
by: Wiederhold, Noah, et al.
Published: (2023)
by: Wiederhold, Noah, et al.
Published: (2023)
Imitating What Works: Simulation-Filtered Modular Policy Learning from Human Videos
by: Zhai, Albert J., et al.
Published: (2026)
by: Zhai, Albert J., et al.
Published: (2026)
FUNCTO: Function-Centric One-Shot Imitation Learning for Tool Manipulation
by: Tang, Chao, et al.
Published: (2025)
by: Tang, Chao, et al.
Published: (2025)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
by: Chen, Zerui, et al.
Published: (2024)
by: Chen, Zerui, et al.
Published: (2024)
2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human Videos
by: Heidinger, Marvin, et al.
Published: (2025)
by: Heidinger, Marvin, et al.
Published: (2025)
AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
by: Gong, Litian, et al.
Published: (2025)
by: Gong, Litian, et al.
Published: (2025)
Cosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modeling
by: He, Yufan, et al.
Published: (2025)
by: He, Yufan, et al.
Published: (2025)
Certified Human Trajectory Prediction
by: Bahari, Mohammadhossein, et al.
Published: (2024)
by: Bahari, Mohammadhossein, et al.
Published: (2024)
Similar Items
-
Rethinking Video Generation Model for the Embodied World
by: Deng, Yufan, et al.
Published: (2026) -
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
by: Fu, Yiyang, et al.
Published: (2026) -
Robot Learning from Human Videos: A Survey
by: Ma, Junyi, et al.
Published: (2026) -
You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations
by: Zhou, Huayi, et al.
Published: (2025) -
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)