Confidence Calibration in Vision-Language-Action Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zollo, Thomas P, Zemel, Richard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation
by: Zollo, Thomas, et al.
Published: (2026)
by: Zollo, Thomas, et al.
Published: (2026)
Test-Time Warmup for Multimodal Large Language Models
by: Rajaneesh, Nikita, et al.
Published: (2025)
by: Rajaneesh, Nikita, et al.
Published: (2025)
Tell Me What To Learn: Generalizing Neural Memory to be Controllable in Natural Language
by: Bennett, Max S., et al.
Published: (2026)
by: Bennett, Max S., et al.
Published: (2026)
Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models
by: Francis-Meretzki, Shelly, et al.
Published: (2026)
by: Francis-Meretzki, Shelly, et al.
Published: (2026)
Adaptive Elicitation of Latent Information Using Natural Language
by: Wang, Jimmy, et al.
Published: (2025)
by: Wang, Jimmy, et al.
Published: (2025)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
OpenVLA: An Open-Source Vision-Language-Action Model
by: Kim, Moo Jin, et al.
Published: (2024)
by: Kim, Moo Jin, et al.
Published: (2024)
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
by: Li, Runze, et al.
Published: (2026)
by: Li, Runze, et al.
Published: (2026)
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
Sigma: The Key for Vision-Language-Action Models toward Telepathic Alignment
by: Wang, Libo
Published: (2025)
by: Wang, Libo
Published: (2025)
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
by: Torne, Marcel, et al.
Published: (2026)
by: Torne, Marcel, et al.
Published: (2026)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
by: Xu, Charles, et al.
Published: (2026)
by: Xu, Charles, et al.
Published: (2026)
Guiding LLM Decision-Making with Fairness Reward Models
by: Hall, Zara, et al.
Published: (2025)
by: Hall, Zara, et al.
Published: (2025)
Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models
by: Zhang, Zhilong, et al.
Published: (2026)
by: Zhang, Zhilong, et al.
Published: (2026)
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
by: Chen, Xiaoyu, et al.
Published: (2025)
by: Chen, Xiaoyu, et al.
Published: (2025)
Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
by: Guan, Weifan, et al.
Published: (2025)
by: Guan, Weifan, et al.
Published: (2025)
OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation
by: Hirose, Noriaki, et al.
Published: (2025)
by: Hirose, Noriaki, et al.
Published: (2025)
Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
by: Lyu, Mingyang, et al.
Published: (2025)
by: Lyu, Mingyang, et al.
Published: (2025)
APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model
by: Lu, Yuanjie, et al.
Published: (2026)
by: Lu, Yuanjie, et al.
Published: (2026)
$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
by: Intelligence, Physical, et al.
Published: (2025)
by: Intelligence, Physical, et al.
Published: (2025)
VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
by: Bi, Jianxin, et al.
Published: (2025)
by: Bi, Jianxin, et al.
Published: (2025)
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
by: Huang, Jialei, et al.
Published: (2025)
by: Huang, Jialei, et al.
Published: (2025)
$π_0$: A Vision-Language-Action Flow Model for General Robot Control
by: Black, Kevin, et al.
Published: (2024)
by: Black, Kevin, et al.
Published: (2024)
Few-Shot Vision-Language Action-Incremental Policy Learning
by: Song, Mingchen, et al.
Published: (2025)
by: Song, Mingchen, et al.
Published: (2025)
Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
by: Driess, Danny, et al.
Published: (2025)
by: Driess, Danny, et al.
Published: (2025)
Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning
by: Hu, Jiaheng, et al.
Published: (2026)
by: Hu, Jiaheng, et al.
Published: (2026)
Run-time Observation Interventions Make Vision-Language-Action Models More Visually Robust
by: Hancock, Asher J., et al.
Published: (2024)
by: Hancock, Asher J., et al.
Published: (2024)
Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
by: Cai, Rui, et al.
Published: (2026)
by: Cai, Rui, et al.
Published: (2026)
From Inference Efficiency to Embodied Efficiency: Revisiting Efficiency Metrics for Vision-Language-Action Models
by: Li, Zhuofan, et al.
Published: (2026)
by: Li, Zhuofan, et al.
Published: (2026)
Habilis-$β$: A Fast-Motion and Long-Lasting On-Device Vision-Language-Action Model
by: Robotics, Tommoro, et al.
Published: (2026)
by: Robotics, Tommoro, et al.
Published: (2026)
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
QuEst: Enhancing Estimates of Quantile-Based Distributional Measures Using Model Predictions
by: Deng, Zhun, et al.
Published: (2025)
by: Deng, Zhun, et al.
Published: (2025)
Conformalized Teleoperation: Confidently Mapping Human Inputs to High-Dimensional Robot Actions
by: Zhao, Michelle, et al.
Published: (2024)
by: Zhao, Michelle, et al.
Published: (2024)
Understanding Asynchronous Inference Methods for Vision-Language-Action Models
by: Agouzoul, Ayoub
Published: (2026)
by: Agouzoul, Ayoub
Published: (2026)
Bridging Language, Vision and Action: Multimodal VAEs in Robotic Manipulation Tasks
by: Sejnova, Gabriela, et al.
Published: (2024)
by: Sejnova, Gabriela, et al.
Published: (2024)
Continuous Reasoning for Vision-Language-Action
by: Wu, Yueh-Hua, et al.
Published: (2026)
by: Wu, Yueh-Hua, et al.
Published: (2026)
TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
by: Im, Hokyun, et al.
Published: (2025)
by: Im, Hokyun, et al.
Published: (2025)
Distribution-Free Statistical Dispersion Control for Societal Applications
by: Deng, Zhun, et al.
Published: (2023)
by: Deng, Zhun, et al.
Published: (2023)
CO-RFT: Efficient Fine-Tuning of Vision-Language-Action Models through Chunked Offline Reinforcement Learning
by: Huang, Dongchi, et al.
Published: (2025)
by: Huang, Dongchi, et al.
Published: (2025)
Similar Items
-
Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation
by: Zollo, Thomas, et al.
Published: (2026) -
Test-Time Warmup for Multimodal Large Language Models
by: Rajaneesh, Nikita, et al.
Published: (2025) -
Tell Me What To Learn: Generalizing Neural Memory to be Controllable in Natural Language
by: Bennett, Max S., et al.
Published: (2026) -
Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models
by: Francis-Meretzki, Shelly, et al.
Published: (2026) -
Adaptive Elicitation of Latent Information Using Natural Language
by: Wang, Jimmy, et al.
Published: (2025)