Confusion-Aware In-Context-Learning for Vision-Language Models in Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | He, Yayun, Kang, Zuheng, Zhao, Botao, Wu, Zhouyin, Peng, Junqing, Wang, Jianzong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization
by: Wang, Jianzong, et al.
Published: (2026)
by: Wang, Jianzong, et al.
Published: (2026)
VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
by: Liu, Chuhang, et al.
Published: (2026)
by: Liu, Chuhang, et al.
Published: (2026)
EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition
by: Li, Pengcheng, et al.
Published: (2025)
by: Li, Pengcheng, et al.
Published: (2025)
Efficient Multi-Model Fusion with Adversarial Complementary Representation Learning
by: Kang, Zuheng, et al.
Published: (2024)
by: Kang, Zuheng, et al.
Published: (2024)
Retrieval-Augmented Audio Deepfake Detection
by: Kang, Zuheng, et al.
Published: (2024)
by: Kang, Zuheng, et al.
Published: (2024)
Generalized Audio Deepfake Detection Using Frame-level Latent Information Entropy
by: Zhao, Botao, et al.
Published: (2025)
by: Zhao, Botao, et al.
Published: (2025)
M3ET: Efficient Vision-Language Learning for Robotics based on Multimodal Mamba-Enhanced Transformer
by: Zhang, Yanxin, et al.
Published: (2025)
by: Zhang, Yanxin, et al.
Published: (2025)
Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
by: Yang, Qingran, et al.
Published: (2026)
by: Yang, Qingran, et al.
Published: (2026)
ACCon: Angle-Compensated Contrastive Regularizer for Deep Regression
by: Zhao, Botao, et al.
Published: (2025)
by: Zhao, Botao, et al.
Published: (2025)
LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
by: Kang, Zeyi, et al.
Published: (2025)
by: Kang, Zeyi, et al.
Published: (2025)
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
by: Yu, Wenda, et al.
Published: (2026)
by: Yu, Wenda, et al.
Published: (2026)
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation
by: Zhao, Wentao, et al.
Published: (2024)
by: Zhao, Wentao, et al.
Published: (2024)
Friction-Aware Safety Locomotion for Wheeled-legged Robots using Vision Language Models and Reinforcement Learning
by: Peng, Bo, et al.
Published: (2024)
by: Peng, Bo, et al.
Published: (2024)
SG-CADVLM: A Context-Aware Decoding Powered Vision Language Model for Safety-Critical Scenario Generation
by: Zhao, Hongyi, et al.
Published: (2026)
by: Zhao, Hongyi, et al.
Published: (2026)
CubeRobot: Grounding Language in Rubik's Cube Manipulation via Vision-Language Model
by: Wang, Feiyang, et al.
Published: (2025)
by: Wang, Feiyang, et al.
Published: (2025)
Task-oriented Robotic Manipulation with Vision Language Models
by: Guran, Nurhan Bulus, et al.
Published: (2024)
by: Guran, Nurhan Bulus, et al.
Published: (2024)
Online Robot Navigation and Manipulation with Distilled Vision-Language Models
by: Liu, Kangcheng
Published: (2024)
by: Liu, Kangcheng
Published: (2024)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
by: Li, Runhao, et al.
Published: (2025)
by: Li, Runhao, et al.
Published: (2025)
VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation
by: Zhang, Chaofan, et al.
Published: (2025)
by: Zhang, Chaofan, et al.
Published: (2025)
OmniVIC: A Self-Improving Variable Impedance Controller with Vision-Language In-Context Learning for Safe Robotic Manipulation
by: Zhang, Heng, et al.
Published: (2025)
by: Zhang, Heng, et al.
Published: (2025)
VIP: Vision Instructed Pre-training for Robotic Manipulation
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
by: Zhao, Enyu, et al.
Published: (2025)
by: Zhao, Enyu, et al.
Published: (2025)
Improving Robotic Manipulation with Efficient Geometry-Aware Vision Encoder
by: Vuong, An Dinh, et al.
Published: (2025)
by: Vuong, An Dinh, et al.
Published: (2025)
SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
by: Fan, Yiguo, et al.
Published: (2025)
by: Fan, Yiguo, et al.
Published: (2025)
Visual Robotic Manipulation with Depth-Aware Pretraining
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation
by: Wang, Zhijie, et al.
Published: (2024)
by: Wang, Zhijie, et al.
Published: (2024)
ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023)
by: Gao, Jensen, et al.
Published: (2023)
CABTO: Context-Aware Behavior Tree Grounding for Robot Manipulation
by: Cai, Yishuai, et al.
Published: (2026)
by: Cai, Yishuai, et al.
Published: (2026)
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
by: Dai, Tingjun, et al.
Published: (2026)
by: Dai, Tingjun, et al.
Published: (2026)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
by: Peng, Xiongfeng, et al.
Published: (2026)
by: Peng, Xiongfeng, et al.
Published: (2026)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
SARO: Space-Aware Robot System for Terrain Crossing via Vision-Language Model
by: Zhu, Shaoting, et al.
Published: (2024)
by: Zhu, Shaoting, et al.
Published: (2024)
FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
by: Miao, Cui, et al.
Published: (2025)
by: Miao, Cui, et al.
Published: (2025)
SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
by: Chen, Qianzhong, et al.
Published: (2025)
by: Chen, Qianzhong, et al.
Published: (2025)
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
by: Li, Pengteng, et al.
Published: (2026)
by: Li, Pengteng, et al.
Published: (2026)
Experiences from Benchmarking Vision-Language-Action Models for Robotic Manipulation
by: Zhang, Yihao, et al.
Published: (2025)
by: Zhang, Yihao, et al.
Published: (2025)
Similar Items
-
Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization
by: Wang, Jianzong, et al.
Published: (2026) -
VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
by: Liu, Chuhang, et al.
Published: (2026) -
EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition
by: Li, Pengcheng, et al.
Published: (2025) -
Efficient Multi-Model Fusion with Adversarial Complementary Representation Learning
by: Kang, Zuheng, et al.
Published: (2024) -
Retrieval-Augmented Audio Deepfake Detection
by: Kang, Zuheng, et al.
Published: (2024)