PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Yupeng, Li, Xiang, Gu, Songen, Zheng, Yuhang, Tian, Shuai, Li, Weize, Wang, Linbo, Fei, Senyu, Li, Pengfei, Gao, Yinfeng, Xing, Zebin, Chen, Yilun, Zhang, Qichao, Li, Haoran, Ding, Wenchao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
by: Gu, Songen, et al.
Published: (2026)
by: Gu, Songen, et al.
Published: (2026)
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
by: Zheng, Yuhang, et al.
Published: (2026)
by: Zheng, Yuhang, et al.
Published: (2026)
UniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set Articulation
by: Jin, Bu, et al.
Published: (2025)
by: Jin, Bu, et al.
Published: (2025)
World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
by: Zheng, Yupeng, et al.
Published: (2025)
by: Zheng, Yupeng, et al.
Published: (2025)
Mimir: Hierarchical Goal-Driven Diffusion with Uncertainty Propagation for End-to-End Autonomous Driving
by: Xing, Zebin, et al.
Published: (2025)
by: Xing, Zebin, et al.
Published: (2025)
Enhancing Indoor Occupancy Prediction via Sparse Query-Based Multi-Level Consistent Knowledge Distillation
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Learning from Mistakes: Post-Training for Driving VLA with Takeover Data
by: Gao, Yinfeng, et al.
Published: (2026)
by: Gao, Yinfeng, et al.
Published: (2026)
Learning High-Frequency Continuous Action Chunks in Latent Space
by: Wang, Kunyun, et al.
Published: (2026)
by: Wang, Kunyun, et al.
Published: (2026)
PlanAgent: A Multi-modal Large Language Agent for Closed-loop Vehicle Motion Planning
by: Zheng, Yupeng, et al.
Published: (2024)
by: Zheng, Yupeng, et al.
Published: (2024)
MonoOcc: Digging into Monocular Semantic Occupancy Prediction
by: Zheng, Yupeng, et al.
Published: (2024)
by: Zheng, Yupeng, et al.
Published: (2024)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
by: Wang, Linbo, et al.
Published: (2026)
by: Wang, Linbo, et al.
Published: (2026)
Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning
by: Liu, Yuan, et al.
Published: (2026)
by: Liu, Yuan, et al.
Published: (2026)
Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
TrajMoE: Scene-Adaptive Trajectory Planning with Mixture of Experts and Reinforcement Learning
by: Xing, Zebin, et al.
Published: (2025)
by: Xing, Zebin, et al.
Published: (2025)
World In Your Hands: A Large-Scale and Open-Source Ecosystem for Learning Human-Centric Manipulation in the Wild
by: Zheng, Yupeng, et al.
Published: (2025)
by: Zheng, Yupeng, et al.
Published: (2025)
NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models
by: Zhu, Ziyue, et al.
Published: (2026)
by: Zhu, Ziyue, et al.
Published: (2026)
Data Scaling Laws for Imitation Learning-Based End-to-End Autonomous Driving
by: Zheng, Yupeng, et al.
Published: (2024)
by: Zheng, Yupeng, et al.
Published: (2024)
TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
by: Jin, Bu, et al.
Published: (2024)
by: Jin, Bu, et al.
Published: (2024)
Generalizable Audio-Visual Navigation via Binaural Difference Attention and Action Transition Prediction
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
by: Li, Yixuan, et al.
Published: (2025)
by: Li, Yixuan, et al.
Published: (2025)
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
by: Xu, Xiaoxu, et al.
Published: (2026)
by: Xu, Xiaoxu, et al.
Published: (2026)
Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement Learning
by: Sun, Jingbo, et al.
Published: (2026)
by: Sun, Jingbo, et al.
Published: (2026)
DySL-VLA: Efficient Vision-Language-Action Model Inference via Dynamic-Static Layer-Skipping for Robot Manipulation
by: Yang, Zebin, et al.
Published: (2026)
by: Yang, Zebin, et al.
Published: (2026)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
by: Xu, Zonghuan, et al.
Published: (2025)
by: Xu, Zonghuan, et al.
Published: (2025)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
by: Li, Boyu, et al.
Published: (2026)
by: Li, Boyu, et al.
Published: (2026)
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
by: Chen, Jiayi, et al.
Published: (2026)
by: Chen, Jiayi, et al.
Published: (2026)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
by: Li, Jiayu, et al.
Published: (2025)
by: Li, Jiayu, et al.
Published: (2025)
PIDNet: Progressive Implicit Decouple Network for Multimodal Action Quality Assessment
by: Li, Qiqi, et al.
Published: (2026)
by: Li, Qiqi, et al.
Published: (2026)
TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control
by: Sun, Yuteng, et al.
Published: (2026)
by: Sun, Yuteng, et al.
Published: (2026)
EaqVLA: Encoding-aligned Quantization for Vision-Language-Action Models
by: Jiang, Feng, et al.
Published: (2025)
by: Jiang, Feng, et al.
Published: (2025)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
by: Zheng, Jinliang, et al.
Published: (2025)
by: Zheng, Jinliang, et al.
Published: (2025)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
by: Xiong, Zheng, et al.
Published: (2025)
by: Xiong, Zheng, et al.
Published: (2025)
Emulating extended Lyman-alpha haloes around star-forming galaxies
by: Li, Pengfei, et al.
Published: (2025)
by: Li, Pengfei, et al.
Published: (2025)
RoboNurse-VLA: Robotic Scrub Nurse System based on Vision-Language-Action Model
by: Li, Shunlei, et al.
Published: (2024)
by: Li, Shunlei, et al.
Published: (2024)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
GaussianGrasper: 3D Language Gaussian Splatting for Open-vocabulary Robotic Grasping
by: Zheng, Yuhang, et al.
Published: (2024)
by: Zheng, Yuhang, et al.
Published: (2024)
LiloDriver: A Lifelong Learning Framework for Closed-loop Motion Planning in Long-tail Autonomous Driving Scenarios
by: Yao, Huaiyuan, et al.
Published: (2025)
by: Yao, Huaiyuan, et al.
Published: (2025)
Similar Items
-
VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
by: Gu, Songen, et al.
Published: (2026) -
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
by: Zheng, Yuhang, et al.
Published: (2026) -
UniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set Articulation
by: Jin, Bu, et al.
Published: (2025) -
World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
by: Zheng, Yupeng, et al.
Published: (2025) -
Mimir: Hierarchical Goal-Driven Diffusion with Uncertainty Propagation for End-to-End Autonomous Driving
by: Xing, Zebin, et al.
Published: (2025)