Gespeichert in:
| Hauptverfasser: | Tur, Yalcin, Naghiyev, Jalal, Fang, Haoquan, Tsai, Wei-Chuan, Duan, Jiafei, Fox, Dieter, Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.07845 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
von: Lin, Zijun, et al.
Veröffentlicht: (2025)
von: Lin, Zijun, et al.
Veröffentlicht: (2025)
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
von: Fang, Haoquan, et al.
Veröffentlicht: (2025)
von: Fang, Haoquan, et al.
Veröffentlicht: (2025)
EVE: Enabling Anyone to Train Robots using Augmented Reality
von: Wang, Jun, et al.
Veröffentlicht: (2024)
von: Wang, Jun, et al.
Veröffentlicht: (2024)
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
von: Pumacay, Wilbert, et al.
Veröffentlicht: (2024)
von: Pumacay, Wilbert, et al.
Veröffentlicht: (2024)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
MolmoAct: Action Reasoning Models that can Reason in Space
von: Lee, Jason, et al.
Veröffentlicht: (2025)
von: Lee, Jason, et al.
Veröffentlicht: (2025)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models
von: Choi, Suhwan, et al.
Veröffentlicht: (2026)
von: Choi, Suhwan, et al.
Veröffentlicht: (2026)
MolmoAct2: Action Reasoning Models for Real-world Deployment
von: Fang, Haoquan, et al.
Veröffentlicht: (2026)
von: Fang, Haoquan, et al.
Veröffentlicht: (2026)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026)
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026)
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025)
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025)
Iterated Learning Improves Compositionality in Large Vision-Language Models
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
von: Chen, Shirui, et al.
Veröffentlicht: (2026)
von: Chen, Shirui, et al.
Veröffentlicht: (2026)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023)
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023)
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
von: Cheng, Long, et al.
Veröffentlicht: (2025)
von: Cheng, Long, et al.
Veröffentlicht: (2025)
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
von: Singh, Ishika, et al.
Veröffentlicht: (2025)
von: Singh, Ishika, et al.
Veröffentlicht: (2025)
Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
von: Rodkin, Ivan, et al.
Veröffentlicht: (2025)
von: Rodkin, Ivan, et al.
Veröffentlicht: (2025)
Latent Transfer Attack: Adversarial Examples via Generative Latent Spaces
von: Shaar, Eitan, et al.
Veröffentlicht: (2026)
von: Shaar, Eitan, et al.
Veröffentlicht: (2026)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
von: Kamath, Amita, et al.
Veröffentlicht: (2026)
von: Kamath, Amita, et al.
Veröffentlicht: (2026)
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
von: Nepal, Aadim, et al.
Veröffentlicht: (2025)
von: Nepal, Aadim, et al.
Veröffentlicht: (2025)
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
von: Kohli, Harsh, et al.
Veröffentlicht: (2026)
von: Kohli, Harsh, et al.
Veröffentlicht: (2026)
GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
von: Abouzeid, Ali, et al.
Veröffentlicht: (2025)
von: Abouzeid, Ali, et al.
Veröffentlicht: (2025)
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
von: Shen, Boyang, et al.
Veröffentlicht: (2026)
von: Shen, Boyang, et al.
Veröffentlicht: (2026)
Two-Scale Latent Dynamics for Recurrent-Depth Transformers
von: Pappone, Francesco, et al.
Veröffentlicht: (2025)
von: Pappone, Francesco, et al.
Veröffentlicht: (2025)
GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation
von: Deshpande, Abhay, et al.
Veröffentlicht: (2025)
von: Deshpande, Abhay, et al.
Veröffentlicht: (2025)
QuoVLA: Quotient Space for Vision-Language-Action Models
von: Wang, Xuan, et al.
Veröffentlicht: (2026)
von: Wang, Xuan, et al.
Veröffentlicht: (2026)
WFM: 3D Wavelet Flow Matching for Ultrafast Multi-Modal MRI Synthesis
von: Tur, Yalcin, et al.
Veröffentlicht: (2026)
von: Tur, Yalcin, et al.
Veröffentlicht: (2026)
OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
von: Liu, Ruixun, et al.
Veröffentlicht: (2025)
von: Liu, Ruixun, et al.
Veröffentlicht: (2025)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
von: Sun, Jingwen, et al.
Veröffentlicht: (2026)
von: Sun, Jingwen, et al.
Veröffentlicht: (2026)
VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model
von: Li, Wenhao, et al.
Veröffentlicht: (2026)
von: Li, Wenhao, et al.
Veröffentlicht: (2026)
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
von: Li, Yixuan, et al.
Veröffentlicht: (2025)
von: Li, Yixuan, et al.
Veröffentlicht: (2025)
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
von: Ye, Angen, et al.
Veröffentlicht: (2025)
von: Ye, Angen, et al.
Veröffentlicht: (2025)
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
von: Wang, Yi Ru, et al.
Veröffentlicht: (2025)
von: Wang, Yi Ru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
von: Lin, Zijun, et al.
Veröffentlicht: (2025) -
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
von: Fang, Haoquan, et al.
Veröffentlicht: (2025) -
EVE: Enabling Anyone to Train Robots using Augmented Reality
von: Wang, Jun, et al.
Veröffentlicht: (2024) -
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
von: Pumacay, Wilbert, et al.
Veröffentlicht: (2024) -
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)