Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Xing, Youguang, Luo, Xu, Xie, Junlin, Gao, Lianli, Shen, Hengtao, Song, Jingkuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
by: Chen, Cheng, et al.
Published: (2024)
by: Chen, Cheng, et al.
Published: (2024)
From Channel Bias to Feature Redundancy: Uncovering the "Less is More" Principle in Few-Shot Learning
by: Zhang, Ji, et al.
Published: (2023)
by: Zhang, Ji, et al.
Published: (2023)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
by: Yin, Zhenhan, et al.
Published: (2025)
by: Yin, Zhenhan, et al.
Published: (2025)
Text-Video Retrieval with Global-Local Semantic Consistent Learning
by: Zhang, Haonan, et al.
Published: (2024)
by: Zhang, Haonan, et al.
Published: (2024)
Policy Contrastive Decoding for Robotic Foundation Models
by: Wu, Shihan, et al.
Published: (2025)
by: Wu, Shihan, et al.
Published: (2025)
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
by: Hu, Yucheng, et al.
Published: (2024)
by: Hu, Yucheng, et al.
Published: (2024)
Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach
by: Yin, Xiaoran, et al.
Published: (2025)
by: Yin, Xiaoran, et al.
Published: (2025)
Benchmarking Few-shot Transferability of Pre-trained Models with Improved Evaluation Protocols
by: Luo, Xu, et al.
Published: (2026)
by: Luo, Xu, et al.
Published: (2026)
AICL: Action In-Context Learning for Video Diffusion Model
by: Liu, Jianzhi, et al.
Published: (2024)
by: Liu, Jianzhi, et al.
Published: (2024)
Turning Video Models into Generalist Robot Policies
by: Li, Sizhe Lester, et al.
Published: (2026)
by: Li, Sizhe Lester, et al.
Published: (2026)
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2026)
by: Van Vo, Tuan, et al.
Published: (2026)
The One RING: a Robotic Indoor Navigation Generalist
by: Eftekhar, Ainaz, et al.
Published: (2024)
by: Eftekhar, Ainaz, et al.
Published: (2024)
Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models
by: Wang, Xuanhan, et al.
Published: (2025)
by: Wang, Xuanhan, et al.
Published: (2025)
F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video Synthesis
by: Su, Sitong, et al.
Published: (2023)
by: Su, Sitong, et al.
Published: (2023)
ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
by: Sun, Yu, et al.
Published: (2026)
by: Sun, Yu, et al.
Published: (2026)
Grounding Bodily Awareness in Visual Representations for Efficient Policy Learning
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
by: Song, Wenxuan, et al.
Published: (2024)
by: Song, Wenxuan, et al.
Published: (2024)
Reliable Few-shot Learning under Dual Noises
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
by: Hou, Zhi, et al.
Published: (2025)
by: Hou, Zhi, et al.
Published: (2025)
What Matters in Building Vision-Language-Action Models for Generalist Robots
by: Li, Xinghang, et al.
Published: (2024)
by: Li, Xinghang, et al.
Published: (2024)
DePT: Decoupled Prompt Tuning
by: Zhang, Ji, et al.
Published: (2023)
by: Zhang, Ji, et al.
Published: (2023)
Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
by: Lyu, Xinyu, et al.
Published: (2024)
by: Lyu, Xinyu, et al.
Published: (2024)
Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation
by: Chen, Beitao, et al.
Published: (2025)
by: Chen, Beitao, et al.
Published: (2025)
CFReID: Continual Few-shot Person Re-Identification
by: Ni, Hao, et al.
Published: (2025)
by: Ni, Hao, et al.
Published: (2025)
A Survey on Efficient Vision-Language-Action Models
by: Yu, Zhaoshu, et al.
Published: (2025)
by: Yu, Zhaoshu, et al.
Published: (2025)
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
by: Liu, Ke, et al.
Published: (2025)
by: Liu, Ke, et al.
Published: (2025)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
by: Chen, Xinyi, et al.
Published: (2025)
by: Chen, Xinyi, et al.
Published: (2025)
The Role of Predictive Uncertainty and Diversity in Embodied AI and Robot Learning
by: Senanayake, Ransalu
Published: (2024)
by: Senanayake, Ransalu
Published: (2024)
Structureless VIO
by: Song, Junlin, et al.
Published: (2025)
by: Song, Junlin, et al.
Published: (2025)
Causal World Modeling for Robot Control
by: Li, Lin, et al.
Published: (2026)
by: Li, Lin, et al.
Published: (2026)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
by: Xie, Senwei, et al.
Published: (2025)
by: Xie, Senwei, et al.
Published: (2025)
EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
by: Punamiya, Ryan, et al.
Published: (2026)
by: Punamiya, Ryan, et al.
Published: (2026)
Cosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modeling
by: He, Yufan, et al.
Published: (2025)
by: He, Yufan, et al.
Published: (2025)
Language-Based Augmentation to Address Shortcut Learning in Object Goal Navigation
by: Hoftijzer, Dennis, et al.
Published: (2024)
by: Hoftijzer, Dennis, et al.
Published: (2024)
Training-Free Semantic Video Composition via Pre-trained Diffusion Model
by: Guo, Jiaqi, et al.
Published: (2024)
by: Guo, Jiaqi, et al.
Published: (2024)
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
by: Yan, Haodong, et al.
Published: (2026)
by: Yan, Haodong, et al.
Published: (2026)
Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
by: Wang, Xuanhan, et al.
Published: (2025)
by: Wang, Xuanhan, et al.
Published: (2025)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
by: Ye, Wencheng, et al.
Published: (2025)
by: Ye, Wencheng, et al.
Published: (2025)
An RGB-D Image Dataset for Lychee Detection and Maturity Classification for Robotic Harvesting
by: Zhang, Zhenpeng, et al.
Published: (2025)
by: Zhang, Zhenpeng, et al.
Published: (2025)
Similar Items
-
CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
by: Chen, Cheng, et al.
Published: (2024) -
From Channel Bias to Feature Redundancy: Uncovering the "Less is More" Principle in Few-Shot Learning
by: Zhang, Ji, et al.
Published: (2023) -
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
by: Yin, Zhenhan, et al.
Published: (2025) -
Text-Video Retrieval with Global-Local Semantic Consistent Learning
by: Zhang, Haonan, et al.
Published: (2024) -
Policy Contrastive Decoding for Robotic Foundation Models
by: Wu, Shihan, et al.
Published: (2025)