UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Jiabing, Chen, Yixiang, Xu, Yuan, Li, Peiyan, Wu, Xiangnan, Wen, Zichen, Fang, Bowen, Yu, Tao, Zhang, Zhengbo, Li, Yingda, Wang, Kai, Liu, Jing, Liu, Nianfeng, Huang, Yan, Wang, Liang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks
von: Chen, Yixiang, et al.
Veröffentlicht: (2026)
von: Chen, Yixiang, et al.
Veröffentlicht: (2026)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
von: Li, Peiyan, et al.
Veröffentlicht: (2026)
von: Li, Peiyan, et al.
Veröffentlicht: (2026)
EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow
von: Chen, Yixiang, et al.
Veröffentlicht: (2025)
von: Chen, Yixiang, et al.
Veröffentlicht: (2025)
LaPA$^2$: Length-Aware Prefix and Prompt Attention Augmentation for Long-Form Controllable Text Generation
von: Yang, Jiabing, et al.
Veröffentlicht: (2025)
von: Yang, Jiabing, et al.
Veröffentlicht: (2025)
EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation
von: Xu, Yuan, et al.
Veröffentlicht: (2025)
von: Xu, Yuan, et al.
Veröffentlicht: (2025)
SKIP: Sparse Keyframe Interpolation Paradigm for Efficient Embodied World Models
von: He, Ziheng, et al.
Veröffentlicht: (2026)
von: He, Ziheng, et al.
Veröffentlicht: (2026)
BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models
von: Li, Peiyan, et al.
Veröffentlicht: (2025)
von: Li, Peiyan, et al.
Veröffentlicht: (2025)
Uncertainty Evaluation of the Caesium Fountain Primary Frequency Standard NIM6
von: Zheng, Fasong, et al.
Veröffentlicht: (2024)
von: Zheng, Fasong, et al.
Veröffentlicht: (2024)
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026)
VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation
von: Chen, Yixiang, et al.
Veröffentlicht: (2025)
von: Chen, Yixiang, et al.
Veröffentlicht: (2025)
Masked Diffusion Vision-Language Models for Temporal Action Localization
von: Wang, Fengshun, et al.
Veröffentlicht: (2026)
von: Wang, Fengshun, et al.
Veröffentlicht: (2026)
Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
Uncertainty-aware Generative Recommendation
von: Fan, Chenxiao, et al.
Veröffentlicht: (2026)
von: Fan, Chenxiao, et al.
Veröffentlicht: (2026)
FloorPlan-VLN: A New Paradigm for Floor Plan Guided Vision-Language Navigation
von: Chen, Kehan, et al.
Veröffentlicht: (2026)
von: Chen, Kehan, et al.
Veröffentlicht: (2026)
What Matters in Building Vision-Language-Action Models for Generalist Robots
von: Li, Xinghang, et al.
Veröffentlicht: (2024)
von: Li, Xinghang, et al.
Veröffentlicht: (2024)
IKOD: Mitigating Visual Attention Degradation in Large Vision-Language Models
von: Yang, Jiabing, et al.
Veröffentlicht: (2025)
von: Yang, Jiabing, et al.
Veröffentlicht: (2025)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
UAU-Net: Uncertainty-aware Representation Learning and Evidential Classification for Facial Action Unit Detection
von: Li, Yuze, et al.
Veröffentlicht: (2026)
von: Li, Yuze, et al.
Veröffentlicht: (2026)
BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents
von: Wang, Shu, et al.
Veröffentlicht: (2025)
von: Wang, Shu, et al.
Veröffentlicht: (2025)
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
von: Yu, Tao, et al.
Veröffentlicht: (2025)
von: Yu, Tao, et al.
Veröffentlicht: (2025)
Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification
von: Huang, Tao, et al.
Veröffentlicht: (2026)
von: Huang, Tao, et al.
Veröffentlicht: (2026)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
Advanced Mathematics Learning Behavior Prediction and Academic Early Warning Model Based on Multimodal Data Analysis
von: Qiong, Liu, et al.
Veröffentlicht: (2026)
von: Qiong, Liu, et al.
Veröffentlicht: (2026)
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
von: Xiong, Minhao, et al.
Veröffentlicht: (2025)
von: Xiong, Minhao, et al.
Veröffentlicht: (2025)
Semi-supervised Node Importance Estimation with Informative Distribution Modeling for Uncertainty Regularization
von: Chen, Yankai, et al.
Veröffentlicht: (2025)
von: Chen, Yankai, et al.
Veröffentlicht: (2025)
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
von: Yang, Yantai, et al.
Veröffentlicht: (2025)
von: Yang, Yantai, et al.
Veröffentlicht: (2025)
Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
ToolWeaver: Weaving Collaborative Semantics for Scalable Tool Use in Large Language Models
von: Fang, Bowen, et al.
Veröffentlicht: (2026)
von: Fang, Bowen, et al.
Veröffentlicht: (2026)
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
von: Wang, Xiaobo, et al.
Veröffentlicht: (2025)
von: Wang, Xiaobo, et al.
Veröffentlicht: (2025)
Ethylene‐Activated E3 Ubiquitin Ligase MdEAEL1 Promotes Apple Fruit Softening by Facilitating the Dissociation of Transcriptional Repressor Complexes
von: Tong Li, et al.
Veröffentlicht: (2025)
von: Tong Li, et al.
Veröffentlicht: (2025)
Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
von: Pei, Xiaohuan, et al.
Veröffentlicht: (2025)
von: Pei, Xiaohuan, et al.
Veröffentlicht: (2025)
FREE: Uncertainty-Aware Autoregression for Parallel Diffusion Transformers
von: Wen, Xinwan, et al.
Veröffentlicht: (2025)
von: Wen, Xinwan, et al.
Veröffentlicht: (2025)
A simulation‐assisted point cloud segmentation neural network for human–robot interaction applications
von: Jingxin Lin, et al.
Veröffentlicht: (2024)
von: Jingxin Lin, et al.
Veröffentlicht: (2024)
Towards Compatible Fine-tuning for Vision-Language Model Updates
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
Data-Driven Revenue Management for Air Cargo
von: Eren, Ezgi, et al.
Veröffentlicht: (2024)
von: Eren, Ezgi, et al.
Veröffentlicht: (2024)
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
Joint Radar Sensing, Location, and Communication Resources Optimization in 6G Network
von: Zhang, Haijun, et al.
Veröffentlicht: (2024)
von: Zhang, Haijun, et al.
Veröffentlicht: (2024)
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
von: Liang, Yuanchang, et al.
Veröffentlicht: (2026)
von: Liang, Yuanchang, et al.
Veröffentlicht: (2026)
MoCha:End-to-End Video Character Replacement without Structural Guidance
von: Xu, Zhengbo, et al.
Veröffentlicht: (2026)
von: Xu, Zhengbo, et al.
Veröffentlicht: (2026)
Supply Chain Finance, Risk Propensity, and Environmental Innovation in Firms: The Moderating Effect of Climate Policy Uncertainty
von: Zhongju Liao, et al.
Veröffentlicht: (2026)
von: Zhongju Liao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks
von: Chen, Yixiang, et al.
Veröffentlicht: (2026) -
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
von: Li, Peiyan, et al.
Veröffentlicht: (2026) -
EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow
von: Chen, Yixiang, et al.
Veröffentlicht: (2025) -
LaPA$^2$: Length-Aware Prefix and Prompt Attention Augmentation for Long-Form Controllable Text Generation
von: Yang, Jiabing, et al.
Veröffentlicht: (2025) -
EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation
von: Xu, Yuan, et al.
Veröffentlicht: (2025)