VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Guanyan, Wang, Meiling, Cui, Te, Mu, Yao, Lu, Haoyang, Zhou, Tianxing, Peng, Zicai, Hu, Mengxiao, Li, Haizhou, Li, Yuan, Yang, Yi, Yue, Yufeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
by: Chen, Guangyan, et al.
Published: (2025)
by: Chen, Guangyan, et al.
Published: (2025)
Human Demonstrations are Generalizable Knowledge for Robots
by: Cui, Te, et al.
Published: (2023)
by: Cui, Te, et al.
Published: (2023)
VIFNet: An End-to-end Visible-Infrared Fusion Network for Image Dehazing
by: Yu, Meng, et al.
Published: (2024)
by: Yu, Meng, et al.
Published: (2024)
TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation Models
by: Yu, Meng, et al.
Published: (2025)
by: Yu, Meng, et al.
Published: (2025)
GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking
by: Chen, Ye, et al.
Published: (2025)
by: Chen, Ye, et al.
Published: (2025)
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
by: Chen, Guangyan, et al.
Published: (2025)
by: Chen, Guangyan, et al.
Published: (2025)
Point Tree Transformer for Point Cloud Registration
by: Wang, Meiling, et al.
Published: (2024)
by: Wang, Meiling, et al.
Published: (2024)
FineSkiing: A Fine-grained Benchmark for Skiing Action Quality Assessment
by: Zhang, Yongji, et al.
Published: (2025)
by: Zhang, Yongji, et al.
Published: (2025)
Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis
by: Jia, Zhenqi, et al.
Published: (2025)
by: Jia, Zhenqi, et al.
Published: (2025)
Primary-Fine Decoupling for Action Generation in Robotic Imitation
by: Lei, Xiaohan, et al.
Published: (2026)
by: Lei, Xiaohan, et al.
Published: (2026)
Towards Fine-grained Interactive Segmentation in Images and Videos
by: Yao, Yuan, et al.
Published: (2025)
by: Yao, Yuan, et al.
Published: (2025)
ICLR: In-Context Imitation Learning with Visual Reasoning
by: Nguyen, Toan, et al.
Published: (2026)
by: Nguyen, Toan, et al.
Published: (2026)
FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification
by: Chen, Xiangyan, et al.
Published: (2025)
by: Chen, Xiangyan, et al.
Published: (2025)
E1 TTS: Simple and Fast Non-Autoregressive TTS
by: Liu, Zhijun, et al.
Published: (2024)
by: Liu, Zhijun, et al.
Published: (2024)
Vision Mamba Distillation for Low-resolution Fine-grained Image Classification
by: Chen, Yao, et al.
Published: (2024)
by: Chen, Yao, et al.
Published: (2024)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment
by: Guo, Geyang, et al.
Published: (2023)
by: Guo, Geyang, et al.
Published: (2023)
SoDA: An Efficient Interaction Paradigm for the Agentic Web
by: Cui, Zicai, et al.
Published: (2025)
by: Cui, Zicai, et al.
Published: (2025)
SkillPager: Query-Adaptive Intra-Skill Navigation via Semantic Node Retrieval
by: Cui, Zicai, et al.
Published: (2026)
by: Cui, Zicai, et al.
Published: (2026)
Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
by: Gu, Difei, et al.
Published: (2025)
by: Gu, Difei, et al.
Published: (2025)
FineParser: A Fine-grained Spatio-temporal Action Parser for Human-centric Action Quality Assessment
by: Xu, Jinglin, et al.
Published: (2024)
by: Xu, Jinglin, et al.
Published: (2024)
FILIC: Dual-Loop Force-Guided Imitation Learning with Impedance Torque Control for Contact-Rich Manipulation Tasks
by: Ge, Haizhou, et al.
Published: (2025)
by: Ge, Haizhou, et al.
Published: (2025)
MA-Bench: Towards Fine-grained Micro-Action Understanding
by: Li, Kun, et al.
Published: (2026)
by: Li, Kun, et al.
Published: (2026)
AVR: Active Vision-Driven Precise Robot Manipulation with Viewpoint and Focal Length Optimization
by: Liu, Yushan, et al.
Published: (2025)
by: Liu, Yushan, et al.
Published: (2025)
A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving
by: Zhang, Liangdong, et al.
Published: (2026)
by: Zhang, Liangdong, et al.
Published: (2026)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
by: Inoue, Sho, et al.
Published: (2024)
by: Inoue, Sho, et al.
Published: (2024)
Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation
by: Chen, Yingjie, et al.
Published: (2025)
by: Chen, Yingjie, et al.
Published: (2025)
GCAM: Gaussian and causal-attention model of food fine-grained recognition
by: Zhuang, Guohang, et al.
Published: (2024)
by: Zhuang, Guohang, et al.
Published: (2024)
Selective Visual Prompting in Vision Mamba
by: Yao, Yifeng, et al.
Published: (2024)
by: Yao, Yifeng, et al.
Published: (2024)
ChartHal: A Fine-grained Framework Evaluating Hallucination of Large Vision Language Models in Chart Understanding
by: Wang, Xingqi, et al.
Published: (2025)
by: Wang, Xingqi, et al.
Published: (2025)
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
by: Zhao, Chen, et al.
Published: (2026)
by: Zhao, Chen, et al.
Published: (2026)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
by: Lin, Yihan, et al.
Published: (2026)
by: Lin, Yihan, et al.
Published: (2026)
SARE: Sample-wise Adaptive Reasoning for Training-free Fine-grained Visual Recognition
by: Yang, Jingxiao, et al.
Published: (2026)
by: Yang, Jingxiao, et al.
Published: (2026)
Fine-grained Image Retrieval via Dual-Vision Adaptation
by: Jiang, Xin, et al.
Published: (2025)
by: Jiang, Xin, et al.
Published: (2025)
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
by: Hait, Soumita, et al.
Published: (2026)
by: Hait, Soumita, et al.
Published: (2026)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
DINO-CoDT: Multi-class Collaborative Detection and Tracking with Vision Foundation Models
by: He, Xunjie, et al.
Published: (2025)
by: He, Xunjie, et al.
Published: (2025)
Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
by: Lyu, Mingyang, et al.
Published: (2025)
by: Lyu, Mingyang, et al.
Published: (2025)
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
by: Hu, Xintong, et al.
Published: (2026)
by: Hu, Xintong, et al.
Published: (2026)
Similar Items
-
FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
by: Chen, Guangyan, et al.
Published: (2025) -
Human Demonstrations are Generalizable Knowledge for Robots
by: Cui, Te, et al.
Published: (2023) -
VIFNet: An End-to-end Visible-Infrared Fusion Network for Image Dehazing
by: Yu, Meng, et al.
Published: (2024) -
TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation Models
by: Yu, Meng, et al.
Published: (2025) -
GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking
by: Chen, Ye, et al.
Published: (2025)