MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Zhenhan, Wang, Xuanhan, Jiang, Jiahao, Deng, Kaiyuan, Chen, Pengqi, Li, Shuangle, Liu, Chong, Xu, Xing, Song, Jingkuan, Gao, Lianli, Shen, Heng Tao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models
by: Wang, Xuanhan, et al.
Published: (2025)
by: Wang, Xuanhan, et al.
Published: (2025)
Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
by: Wang, Xuanhan, et al.
Published: (2025)
by: Wang, Xuanhan, et al.
Published: (2025)
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
by: Liu, Ke, et al.
Published: (2025)
by: Liu, Ke, et al.
Published: (2025)
Sim-and-Human Co-training for Data-Efficient and Generalizable Robotic Manipulation
by: Fang, Kaipeng, et al.
Published: (2026)
by: Fang, Kaipeng, et al.
Published: (2026)
Beyond the Majority: Long-tail Imitation Learning for Robotic Manipulation
by: Zhu, Junhong, et al.
Published: (2026)
by: Zhu, Junhong, et al.
Published: (2026)
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
by: Wu, Shihan, et al.
Published: (2024)
by: Wu, Shihan, et al.
Published: (2024)
Benchmarking Few-shot Transferability of Pre-trained Models with Improved Evaluation Protocols
by: Luo, Xu, et al.
Published: (2026)
by: Luo, Xu, et al.
Published: (2026)
Any Target Can be Offense: Adversarial Example Generation via Generalized Latent Infection
by: Sun, Youheng, et al.
Published: (2024)
by: Sun, Youheng, et al.
Published: (2024)
Pseudo-Label Refinement for Robust Wheat Head Segmentation via Two-Stage Hybrid Training
by: Jiang, Jiahao, et al.
Published: (2025)
by: Jiang, Jiahao, et al.
Published: (2025)
Training-Free Semantic Video Composition via Pre-trained Diffusion Model
by: Guo, Jiaqi, et al.
Published: (2024)
by: Guo, Jiaqi, et al.
Published: (2024)
AICL: Action In-Context Learning for Video Diffusion Model
by: Liu, Jianzhi, et al.
Published: (2024)
by: Liu, Jianzhi, et al.
Published: (2024)
InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
Practical No-box Adversarial Attacks with Training-free Hybrid Image Transformation
by: Zhang, Qilong, et al.
Published: (2022)
by: Zhang, Qilong, et al.
Published: (2022)
Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
by: Lyu, Xinyu, et al.
Published: (2024)
by: Lyu, Xinyu, et al.
Published: (2024)
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
by: Su, Taiyi, et al.
Published: (2026)
by: Su, Taiyi, et al.
Published: (2026)
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
by: Xing, Youguang, et al.
Published: (2025)
by: Xing, Youguang, et al.
Published: (2025)
A Survey on Efficient Vision-Language-Action Models
by: Yu, Zhaoshu, et al.
Published: (2025)
by: Yu, Zhaoshu, et al.
Published: (2025)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
Robotic Imitation of Human Actions
by: Spisak, Josua, et al.
Published: (2024)
by: Spisak, Josua, et al.
Published: (2024)
Policy Contrastive Decoding for Robotic Foundation Models
by: Wu, Shihan, et al.
Published: (2025)
by: Wu, Shihan, et al.
Published: (2025)
ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models
by: Li, Puhao, et al.
Published: (2025)
by: Li, Puhao, et al.
Published: (2025)
Reliable Few-shot Learning under Dual Noises
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
DePT: Decoupled Prompt Tuning
by: Zhang, Ji, et al.
Published: (2023)
by: Zhang, Ji, et al.
Published: (2023)
Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation
by: Chen, Beitao, et al.
Published: (2025)
by: Chen, Beitao, et al.
Published: (2025)
CFReID: Continual Few-shot Person Re-Identification
by: Ni, Hao, et al.
Published: (2025)
by: Ni, Hao, et al.
Published: (2025)
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
by: Chen, Beitao, et al.
Published: (2025)
by: Chen, Beitao, et al.
Published: (2025)
Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach
by: Yin, Xiaoran, et al.
Published: (2025)
by: Yin, Xiaoran, et al.
Published: (2025)
HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model
by: Zhu, Xiang, et al.
Published: (2026)
by: Zhu, Xiang, et al.
Published: (2026)
F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video Synthesis
by: Su, Sitong, et al.
Published: (2023)
by: Su, Sitong, et al.
Published: (2023)
Language-Grounded Decoupled Action Representation for Robotic Manipulation
by: Weng, Wuding, et al.
Published: (2026)
by: Weng, Wuding, et al.
Published: (2026)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
by: Ma, Guoqing, et al.
Published: (2026)
by: Ma, Guoqing, et al.
Published: (2026)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
by: Chen, Yandu, et al.
Published: (2025)
by: Chen, Yandu, et al.
Published: (2025)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
by: Li, Runhao, et al.
Published: (2025)
by: Li, Runhao, et al.
Published: (2025)
GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
by: Deng, Shengliang, et al.
Published: (2025)
by: Deng, Shengliang, et al.
Published: (2025)
From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings
by: Zhang, Jiajie, et al.
Published: (2025)
by: Zhang, Jiajie, et al.
Published: (2025)
Debiased Orthogonal Boundary-Driven Efficient Noise Mitigation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation
by: Kobayashi, Masato, et al.
Published: (2025)
by: Kobayashi, Masato, et al.
Published: (2025)
SeMv-3D: Towards Concurrency of Semantic and Multi-view Consistency in General Text-to-3D Generation
by: Cai, Xiao, et al.
Published: (2024)
by: Cai, Xiao, et al.
Published: (2024)
Similar Items
-
Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models
by: Wang, Xuanhan, et al.
Published: (2025) -
Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
by: Wang, Xuanhan, et al.
Published: (2025) -
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
by: Liu, Ke, et al.
Published: (2025) -
Sim-and-Human Co-training for Data-Efficient and Generalizable Robotic Manipulation
by: Fang, Kaipeng, et al.
Published: (2026) -
Beyond the Majority: Long-tail Imitation Learning for Robotic Manipulation
by: Zhu, Junhong, et al.
Published: (2026)