FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Guangyan, Wang, Meiling, Cui, Te, Mu, Yao, Lu, Haoyang, Peng, Zicai, Hu, Mengxiao, Zhou, Tianxing, Fu, Mengyin, Yang, Yi, Yue, Yufeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
von: Chen, Guanyan, et al.
Veröffentlicht: (2024)
von: Chen, Guanyan, et al.
Veröffentlicht: (2024)
Human Demonstrations are Generalizable Knowledge for Robots
von: Cui, Te, et al.
Veröffentlicht: (2023)
von: Cui, Te, et al.
Veröffentlicht: (2023)
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
von: Chen, Guangyan, et al.
Veröffentlicht: (2025)
von: Chen, Guangyan, et al.
Veröffentlicht: (2025)
Unified Vertex Motion Estimation for Integrated Video Stabilization and Stitching in Tractor-Trailer Wheeled Robots
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
Point Tree Transformer for Point Cloud Registration
von: Wang, Meiling, et al.
Veröffentlicht: (2024)
von: Wang, Meiling, et al.
Veröffentlicht: (2024)
STEP Planner: Constructing cross-hierarchical subgoal tree as an embodied long-horizon task planner
von: Zhou, Tianxing, et al.
Veröffentlicht: (2025)
von: Zhou, Tianxing, et al.
Veröffentlicht: (2025)
OmniMap: A General Mapping Framework Integrating Optics, Geometry, and Semantics
von: Deng, Yinan, et al.
Veröffentlicht: (2025)
von: Deng, Yinan, et al.
Veröffentlicht: (2025)
GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking
von: Chen, Ye, et al.
Veröffentlicht: (2025)
von: Chen, Ye, et al.
Veröffentlicht: (2025)
Identifying Influential Actions in Human-Robot Interactions
von: Jiang, Haoyang, et al.
Veröffentlicht: (2026)
von: Jiang, Haoyang, et al.
Veröffentlicht: (2026)
OpenIN: Open-Vocabulary Instance-Oriented Navigation in Dynamic Domestic Environments
von: Tang, Yujie, et al.
Veröffentlicht: (2025)
von: Tang, Yujie, et al.
Veröffentlicht: (2025)
OpenGraph: Open-Vocabulary Hierarchical 3D Graph Representation in Large-Scale Outdoor Environments
von: Deng, Yinan, et al.
Veröffentlicht: (2024)
von: Deng, Yinan, et al.
Veröffentlicht: (2024)
OpenVox: Real-time Instance-level Open-vocabulary Probabilistic Voxel Representation
von: Deng, Yinan, et al.
Veröffentlicht: (2025)
von: Deng, Yinan, et al.
Veröffentlicht: (2025)
OpenObject-NAV: Open-Vocabulary Object-Oriented Navigation Based on Dynamic Carrier-Relationship Scene Graph
von: Tang, Yujie, et al.
Veröffentlicht: (2024)
von: Tang, Yujie, et al.
Veröffentlicht: (2024)
DexHiL: A Human-in-the-Loop Framework for Vision-Language-Action Model Post-Training in Dexterous Manipulation
von: Han, Yifan, et al.
Veröffentlicht: (2026)
von: Han, Yifan, et al.
Veröffentlicht: (2026)
Human2Robot: Learning Robot Actions from Paired Human-Robot Videos
von: Xie, Sicheng, et al.
Veröffentlicht: (2025)
von: Xie, Sicheng, et al.
Veröffentlicht: (2025)
$π_0$-EqM: Equilibrium Matching for Closed-Loop Vision-Language-Action Control
von: Liu, Huanming, et al.
Veröffentlicht: (2026)
von: Liu, Huanming, et al.
Veröffentlicht: (2026)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
JailWAM: Jailbreaking World Action Models in Robot Control
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
AVR: Active Vision-Driven Precise Robot Manipulation with Viewpoint and Focal Length Optimization
von: Liu, Yushan, et al.
Veröffentlicht: (2025)
von: Liu, Yushan, et al.
Veröffentlicht: (2025)
AINav: Large Language Model-Based Adaptive Interactive Navigation
von: Zhou, Kangjie, et al.
Veröffentlicht: (2025)
von: Zhou, Kangjie, et al.
Veröffentlicht: (2025)
HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos
von: Weng, Haoyang, et al.
Veröffentlicht: (2025)
von: Weng, Haoyang, et al.
Veröffentlicht: (2025)
HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
Temporal Action Selection for Action Chunking
von: Weng, Yueyang, et al.
Veröffentlicht: (2025)
von: Weng, Yueyang, et al.
Veröffentlicht: (2025)
ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
von: Dai, Weisheng, et al.
Veröffentlicht: (2026)
von: Dai, Weisheng, et al.
Veröffentlicht: (2026)
CyboRacket: A Perception-to-Action Framework for Humanoid Racket Sports
von: Ren, Peng, et al.
Veröffentlicht: (2026)
von: Ren, Peng, et al.
Veröffentlicht: (2026)
A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving
von: Zhang, Liangdong, et al.
Veröffentlicht: (2026)
von: Zhang, Liangdong, et al.
Veröffentlicht: (2026)
DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models
von: Jia, Emily Yue-Ting, et al.
Veröffentlicht: (2026)
von: Jia, Emily Yue-Ting, et al.
Veröffentlicht: (2026)
OpenObj: Open-Vocabulary Object-Level Neural Radiance Fields with Fine-Grained Understanding
von: Deng, Yinan, et al.
Veröffentlicht: (2024)
von: Deng, Yinan, et al.
Veröffentlicht: (2024)
STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models
von: Xu, Feng, et al.
Veröffentlicht: (2025)
von: Xu, Feng, et al.
Veröffentlicht: (2025)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
von: Ma, Teli, et al.
Veröffentlicht: (2026)
von: Ma, Teli, et al.
Veröffentlicht: (2026)
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
von: Hu, Xintong, et al.
Veröffentlicht: (2026)
von: Hu, Xintong, et al.
Veröffentlicht: (2026)
TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation Models
von: Yu, Meng, et al.
Veröffentlicht: (2025)
von: Yu, Meng, et al.
Veröffentlicht: (2025)
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version)
von: Mu, Yao, et al.
Veröffentlicht: (2024)
von: Mu, Yao, et al.
Veröffentlicht: (2024)
PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation
von: Yin, Yifan, et al.
Veröffentlicht: (2025)
von: Yin, Yifan, et al.
Veröffentlicht: (2025)
Phoenix: A Motion-based Self-Reflection Framework for Fine-grained Robotic Action Correction
von: Xia, Wenke, et al.
Veröffentlicht: (2025)
von: Xia, Wenke, et al.
Veröffentlicht: (2025)
AttenA+: Rectifying Action Inequality in Robotic Foundation Models
von: Peng, Daojie, et al.
Veröffentlicht: (2026)
von: Peng, Daojie, et al.
Veröffentlicht: (2026)
GrainGrasp: Dexterous Grasp Generation with Fine-grained Contact Guidance
von: Zhao, Fuqiang, et al.
Veröffentlicht: (2024)
von: Zhao, Fuqiang, et al.
Veröffentlicht: (2024)
$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation
von: Zhou, Pengfei, et al.
Veröffentlicht: (2026)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2026)
DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos
von: Mu, Juncheng, et al.
Veröffentlicht: (2026)
von: Mu, Juncheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
von: Chen, Guanyan, et al.
Veröffentlicht: (2024) -
Human Demonstrations are Generalizable Knowledge for Robots
von: Cui, Te, et al.
Veröffentlicht: (2023) -
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
von: Chen, Guangyan, et al.
Veröffentlicht: (2025) -
Unified Vertex Motion Estimation for Integrated Video Stabilization and Stitching in Tractor-Trailer Wheeled Robots
von: Liang, Hao, et al.
Veröffentlicht: (2024) -
Point Tree Transformer for Point Cloud Registration
von: Wang, Meiling, et al.
Veröffentlicht: (2024)