RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kwok, Jacky, Agia, Christopher, Sinha, Rohan, Foutter, Matt, Li, Shulu, Stoica, Ion, Mirhoseini, Azalia, Pavone, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment
by: Kwok, Jacky, et al.
Published: (2026)
by: Kwok, Jacky, et al.
Published: (2026)
Real-Time Anomaly Detection and Reactive Planning with Large Language Models
by: Sinha, Rohan, et al.
Published: (2024)
by: Sinha, Rohan, et al.
Published: (2024)
CodeMonkeys: Scaling Test-Time Compute for Software Engineering
by: Ehrlich, Ryan, et al.
Published: (2025)
by: Ehrlich, Ryan, et al.
Published: (2025)
On the Role of Temperature Sampling in Test-Time Scaling
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
by: Brown, Bradley, et al.
Published: (2024)
by: Brown, Bradley, et al.
Published: (2024)
Real-Time Out-of-Distribution Failure Prevention via Multi-Modal Reasoning
by: Ganai, Milan, et al.
Published: (2025)
by: Ganai, Milan, et al.
Published: (2025)
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
by: Agia, Christopher, et al.
Published: (2024)
by: Agia, Christopher, et al.
Published: (2024)
Preventing Robotic Jailbreaking via Multimodal Domain Adaptation
by: Marchiori, Francesco, et al.
Published: (2025)
by: Marchiori, Francesco, et al.
Published: (2025)
Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models
by: Costello, Caia, et al.
Published: (2025)
by: Costello, Caia, et al.
Published: (2025)
Text2Interaction: Establishing Safe and Preferable Human-Robot Interaction
by: Thumm, Jakob, et al.
Published: (2024)
by: Thumm, Jakob, et al.
Published: (2024)
RoboMorph: In-Context Meta-Learning for Robot Dynamics Modeling
by: Bazzi, Manuel Bianchi, et al.
Published: (2024)
by: Bazzi, Manuel Bianchi, et al.
Published: (2024)
HPRM: High-Performance Robotic Middleware for Intelligent Autonomous Systems
by: Kwok, Jacky, et al.
Published: (2024)
by: Kwok, Jacky, et al.
Published: (2024)
CUPID: Curating Data your Robot Loves with Influence Functions
by: Agia, Christopher, et al.
Published: (2025)
by: Agia, Christopher, et al.
Published: (2025)
Deployment-Time Reliability of Learned Robot Policies
by: Agia, Christopher
Published: (2026)
by: Agia, Christopher
Published: (2026)
That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design
by: Goldie, Anna, et al.
Published: (2024)
by: Goldie, Anna, et al.
Published: (2024)
How Do Large Language Monkeys Get Their Power (Laws)?
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Text2Motion: From Natural Language Instructions to Feasible Plans
by: Lin, Kevin, et al.
Published: (2023)
by: Lin, Kevin, et al.
Published: (2023)
Contextual Graph Representations for Task-Driven 3D Perception and Planning
by: Agia, Christopher
Published: (2026)
by: Agia, Christopher
Published: (2026)
Realistic Extreme Behavior Generation for Improved AV Testing
by: Dyro, Robert, et al.
Published: (2024)
by: Dyro, Robert, et al.
Published: (2024)
Vision Foundation Model Embedding-Based Semantic Anomaly Detection
by: Ronecker, Max Peter, et al.
Published: (2025)
by: Ronecker, Max Peter, et al.
Published: (2025)
Some Present-Day Problems of Romanian Library Science
by: Stoica, Ion
Published: (1973)
by: Stoica, Ion
Published: (1973)
The Central University Library, Bucharest. Over Seventy-five Years in the History of a Collection
by: Stoica, Ion
Published: (1972)
by: Stoica, Ion
Published: (1972)
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
by: Goldie, Anna, et al.
Published: (2025)
by: Goldie, Anna, et al.
Published: (2025)
RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models
by: Kim, Dongyoung, et al.
Published: (2026)
by: Kim, Dongyoung, et al.
Published: (2026)
Federation of Experts: Communication Efficient Distributed Inference for Large Language Models
by: Abdurrahman, Muhammad Shahir, et al.
Published: (2026)
by: Abdurrahman, Muhammad Shahir, et al.
Published: (2026)
TRACE: Capability-Targeted Agentic Training
by: Kang, Hangoo, et al.
Published: (2026)
by: Kang, Hangoo, et al.
Published: (2026)
ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling
by: Winston, Caleb, et al.
Published: (2026)
by: Winston, Caleb, et al.
Published: (2026)
Robo-taxi Fleet Coordination at Scale via Reinforcement Learning
by: Tresca, Luigi, et al.
Published: (2025)
by: Tresca, Luigi, et al.
Published: (2025)
Space-LLaVA: a Vision-Language Model Adapted to Extraterrestrial Applications
by: Foutter, Matthew, et al.
Published: (2024)
by: Foutter, Matthew, et al.
Published: (2024)
S*: Test Time Scaling for Code Generation
by: Li, Dacheng, et al.
Published: (2025)
by: Li, Dacheng, et al.
Published: (2025)
Hydragen: High-Throughput LLM Inference with Shared Prefixes
by: Juravsky, Jordan, et al.
Published: (2024)
by: Juravsky, Jordan, et al.
Published: (2024)
CHESS: Contextual Harnessing for Efficient SQL Synthesis
by: Talaei, Shayan, et al.
Published: (2024)
by: Talaei, Shayan, et al.
Published: (2024)
CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models
by: Lee, Donghyun, et al.
Published: (2024)
by: Lee, Donghyun, et al.
Published: (2024)
$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
by: Cemri, Mert, et al.
Published: (2025)
by: Cemri, Mert, et al.
Published: (2025)
Observing and Controlling Features in Vision-Language-Action Models
by: Buurmeijer, Hugo, et al.
Published: (2026)
by: Buurmeijer, Hugo, et al.
Published: (2026)
SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning Models
by: Biju, Emil, et al.
Published: (2025)
by: Biju, Emil, et al.
Published: (2025)
On Nonlinear Inertial Transformations
by: Agia, Nicholas
Published: (2025)
by: Agia, Nicholas
Published: (2025)
KernelBench: Can LLMs Write Efficient GPU Kernels?
by: Ouyang, Anne, et al.
Published: (2025)
by: Ouyang, Anne, et al.
Published: (2025)
Shrinking the Generation-Verification Gap with Weak Verifiers
by: Saad-Falcon, Jon, et al.
Published: (2025)
by: Saad-Falcon, Jon, et al.
Published: (2025)
Similar Items
-
Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment
by: Kwok, Jacky, et al.
Published: (2026) -
Real-Time Anomaly Detection and Reactive Planning with Large Language Models
by: Sinha, Rohan, et al.
Published: (2024) -
CodeMonkeys: Scaling Test-Time Compute for Software Engineering
by: Ehrlich, Ryan, et al.
Published: (2025) -
On the Role of Temperature Sampling in Test-Time Scaling
by: Wu, Yuheng, et al.
Published: (2025) -
Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
by: Brown, Bradley, et al.
Published: (2024)