Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Chenyou, Yan, Fangzheng, Bai, Chenjia, Wang, Jiepeng, Zhang, Chi, Wang, Zhen, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training
by: He, Haoran, et al.
Published: (2024)
by: He, Haoran, et al.
Published: (2024)
Planning-Guided Diffusion Policy Learning for Generalizable Contact-Rich Bimanual Manipulation
by: Li, Xuanlin, et al.
Published: (2024)
by: Li, Xuanlin, et al.
Published: (2024)
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
by: Zhang, Junjie, et al.
Published: (2024)
by: Zhang, Junjie, et al.
Published: (2024)
Towards Generalizable Robotic Manipulation in Dynamic Environments
by: Fang, Heng, et al.
Published: (2026)
by: Fang, Heng, et al.
Published: (2026)
GraspLDP: Towards Generalizable Grasping Policy via Latent Diffusion
by: Xiang, Enda, et al.
Published: (2026)
by: Xiang, Enda, et al.
Published: (2026)
Efficient Training of Generalizable Visuomotor Policies via Control-Aware Augmentation
by: Zhao, Yinuo, et al.
Published: (2024)
by: Zhao, Yinuo, et al.
Published: (2024)
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
by: Hu, Yucheng, et al.
Published: (2024)
by: Hu, Yucheng, et al.
Published: (2024)
Re$^2$MoGen: Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement
by: Zheng, Jiakun, et al.
Published: (2026)
by: Zheng, Jiakun, et al.
Published: (2026)
You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations
by: Zhou, Huayi, et al.
Published: (2025)
by: Zhou, Huayi, et al.
Published: (2025)
PanoSLAM: Panoptic 3D Scene Reconstruction via Gaussian SLAM
by: Chen, Runnan, et al.
Published: (2024)
by: Chen, Runnan, et al.
Published: (2024)
FoldNet: Learning Generalizable Closed-Loop Policy for Garment Folding via Keypoint-Driven Asset and Demonstration Synthesis
by: Chen, Yuxing, et al.
Published: (2025)
by: Chen, Yuxing, et al.
Published: (2025)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
by: Gong, Zhefei, et al.
Published: (2024)
by: Gong, Zhefei, et al.
Published: (2024)
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
by: Wen, Junjie, et al.
Published: (2024)
by: Wen, Junjie, et al.
Published: (2024)
Adapt2Reward: Adapting Video-Language Models to Generalizable Robotic Rewards via Failure Prompts
by: Yang, Yanting, et al.
Published: (2024)
by: Yang, Yanting, et al.
Published: (2024)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
by: Garcia, Ricardo, et al.
Published: (2024)
by: Garcia, Ricardo, et al.
Published: (2024)
ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning
by: Li, Kailin, et al.
Published: (2025)
by: Li, Kailin, et al.
Published: (2025)
ReSeFlow: Rectifying SE(3)-Equivariant Policy Learning Flows
by: Wang, Zhitao, et al.
Published: (2025)
by: Wang, Zhitao, et al.
Published: (2025)
3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
by: Ze, Yanjie, et al.
Published: (2024)
by: Ze, Yanjie, et al.
Published: (2024)
StructBiHOI: Structured Articulation Modeling for Long--Horizon Bimanual Hand--Object Interaction Generation
by: Wang, Zhi, et al.
Published: (2026)
by: Wang, Zhi, et al.
Published: (2026)
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024)
by: Bharadhwaj, Homanga, et al.
Published: (2024)
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
by: Liu, Songming, et al.
Published: (2024)
by: Liu, Songming, et al.
Published: (2024)
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
by: Chen, Zerui, et al.
Published: (2026)
by: Chen, Zerui, et al.
Published: (2026)
Large Video Planner Enables Generalizable Robot Control
by: Chen, Boyuan, et al.
Published: (2025)
by: Chen, Boyuan, et al.
Published: (2025)
BiFold: Bimanual Cloth Folding with Language Guidance
by: Barbany, Oriol, et al.
Published: (2025)
by: Barbany, Oriol, et al.
Published: (2025)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025)
by: Shen, Yichao, et al.
Published: (2025)
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
by: Fan, Qingyu, et al.
Published: (2026)
by: Fan, Qingyu, et al.
Published: (2026)
FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation
by: Wang, Sen, et al.
Published: (2025)
by: Wang, Sen, et al.
Published: (2025)
DexMan: Learning Bimanual Dexterous Manipulation from Human and Generated Videos
by: Hsieh, Jhen, et al.
Published: (2025)
by: Hsieh, Jhen, et al.
Published: (2025)
ReliOcc: Towards Reliable Semantic Occupancy Prediction via Uncertainty Learning
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy
by: Wu, Pengyuan, et al.
Published: (2026)
by: Wu, Pengyuan, et al.
Published: (2026)
Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning
by: Bai, Chenjia, et al.
Published: (2020)
by: Bai, Chenjia, et al.
Published: (2020)
Generalizable Humanoid Manipulation with 3D Diffusion Policies
by: Ze, Yanjie, et al.
Published: (2024)
by: Ze, Yanjie, et al.
Published: (2024)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
by: Xie, Senwei, et al.
Published: (2025)
by: Xie, Senwei, et al.
Published: (2025)
Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human Videos
by: Heidinger, Marvin, et al.
Published: (2025)
by: Heidinger, Marvin, et al.
Published: (2025)
ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
by: Wu, Yuhan, et al.
Published: (2025)
by: Wu, Yuhan, et al.
Published: (2025)
CRAFT: Video Diffusion for Bimanual Robot Data Generation
by: Chen, Jason, et al.
Published: (2026)
by: Chen, Jason, et al.
Published: (2026)
DexGarmentLab: Dexterous Garment Manipulation Environment with Generalizable Policy
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
NVSPolicy: Adaptive Novel-View Synthesis for Generalizable Language-Conditioned Policy Learning
by: Shi, Le, et al.
Published: (2025)
by: Shi, Le, et al.
Published: (2025)
Similar Items
-
Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training
by: He, Haoran, et al.
Published: (2024) -
Planning-Guided Diffusion Policy Learning for Generalizable Contact-Rich Bimanual Manipulation
by: Li, Xuanlin, et al.
Published: (2024) -
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
by: Zhang, Junjie, et al.
Published: (2024) -
Towards Generalizable Robotic Manipulation in Dynamic Environments
by: Fang, Heng, et al.
Published: (2026) -
GraspLDP: Towards Generalizable Grasping Policy via Latent Diffusion
by: Xiang, Enda, et al.
Published: (2026)