SITCOM: Scaling Inference-Time COMpute for VLAs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saxena, Ayudh, Shah, Harsh, Routray, Sandeep, Shah, Rishi Rajesh, Pahwa, Esha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2025)
von: Tang, Jiaming, et al.
Veröffentlicht: (2025)
FASTER: Rethinking Real-Time Flow VLAs
von: Lu, Yuxiang, et al.
Veröffentlicht: (2026)
von: Lu, Yuxiang, et al.
Veröffentlicht: (2026)
Running VLAs at Real-time Speed
von: Ma, Yunchao, et al.
Veröffentlicht: (2025)
von: Ma, Yunchao, et al.
Veröffentlicht: (2025)
Realtime-VLA FLASH: Speculative Inference Framework for Diffusion-based VLAs
von: Niu, Jiahui, et al.
Veröffentlicht: (2026)
von: Niu, Jiahui, et al.
Veröffentlicht: (2026)
VLAs are Confined yet Capable of Generalizing to Novel Instructions
von: Li, Quanyi
Veröffentlicht: (2025)
von: Li, Quanyi
Veröffentlicht: (2025)
Shallow-π: Knowledge Distillation for Flow-based VLAs
von: Jeon, Boseong, et al.
Veröffentlicht: (2026)
von: Jeon, Boseong, et al.
Veröffentlicht: (2026)
Primitive Subspaces Mediate Few-Shot Transfer in VLAs
von: Singh, Anya, et al.
Veröffentlicht: (2026)
von: Singh, Anya, et al.
Veröffentlicht: (2026)
Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs
von: Zhao, Jianchao, et al.
Veröffentlicht: (2026)
von: Zhao, Jianchao, et al.
Veröffentlicht: (2026)
Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
von: Hancock, Asher J., et al.
Veröffentlicht: (2025)
von: Hancock, Asher J., et al.
Veröffentlicht: (2025)
Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks
von: Haresh, Sanjay, et al.
Veröffentlicht: (2026)
von: Haresh, Sanjay, et al.
Veröffentlicht: (2026)
How Do VLAs Effectively Inherit from VLMs?
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)
cVLA: Towards Efficient Camera-Space VLAs
von: Argus, Max, et al.
Veröffentlicht: (2025)
von: Argus, Max, et al.
Veröffentlicht: (2025)
How VLAs (Really) Work In Open-World Environments
von: Rasouli, Amir, et al.
Veröffentlicht: (2026)
von: Rasouli, Amir, et al.
Veröffentlicht: (2026)
Scaling Sim-to-Real Reinforcement Learning for Robot VLAs with Generative 3D Worlds
von: Choi, Andrew, et al.
Veröffentlicht: (2026)
von: Choi, Andrew, et al.
Veröffentlicht: (2026)
Realtime-VLA V2: Learning to Run VLAs Fast, Smooth, and Accurate
von: Yang, Chen, et al.
Veröffentlicht: (2026)
von: Yang, Chen, et al.
Veröffentlicht: (2026)
VLA-0: Building State-of-the-Art VLAs with Zero Modification
von: Goyal, Ankit, et al.
Veröffentlicht: (2025)
von: Goyal, Ankit, et al.
Veröffentlicht: (2025)
REALM: Real-Time Estimates of Assistance for Learned Models in Human-Robot Interaction
von: Hagenow, Michael, et al.
Veröffentlicht: (2025)
von: Hagenow, Michael, et al.
Veröffentlicht: (2025)
Differentiate-and-Inject: Enhancing VLAs via Functional Differentiation Induced by In-Parameter Structural Reasoning
von: Hou, Jingyi, et al.
Veröffentlicht: (2026)
von: Hou, Jingyi, et al.
Veröffentlicht: (2026)
Do World Action Models Generalize Better than VLAs? A Robustness Study
von: Zhang, Zhanguang, et al.
Veröffentlicht: (2026)
von: Zhang, Zhanguang, et al.
Veröffentlicht: (2026)
Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs
von: Priyadershi, Abhinaw, et al.
Veröffentlicht: (2026)
von: Priyadershi, Abhinaw, et al.
Veröffentlicht: (2026)
Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
AGDC: Automatic Garbage Detection and Collection
von: Bansal, Siddhant, et al.
Veröffentlicht: (2019)
von: Bansal, Siddhant, et al.
Veröffentlicht: (2019)
LEGS: Fine-Tuning Teleop-Free VLAs for Humanoid Loco-manipulation in an Embodied Gaussian Splatting World
von: Kim, Hojune, et al.
Veröffentlicht: (2026)
von: Kim, Hojune, et al.
Veröffentlicht: (2026)
ViPRA: Video Prediction for Robot Actions
von: Routray, Sandeep, et al.
Veröffentlicht: (2025)
von: Routray, Sandeep, et al.
Veröffentlicht: (2025)
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
von: Fang, Yu, et al.
Veröffentlicht: (2026)
von: Fang, Yu, et al.
Veröffentlicht: (2026)
Compositional Visual Planning via Inference-Time Diffusion Scaling
von: Zhang, Yixin, et al.
Veröffentlicht: (2026)
von: Zhang, Yixin, et al.
Veröffentlicht: (2026)
The Price Is Not Right: Neuro-Symbolic Methods Outperform VLAs on Structured Long-Horizon Manipulation Tasks with Significantly Lower Energy Consumption
von: Duggan, Timothy, et al.
Veröffentlicht: (2026)
von: Duggan, Timothy, et al.
Veröffentlicht: (2026)
What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies
von: Zhang, Jiachen, et al.
Veröffentlicht: (2026)
von: Zhang, Jiachen, et al.
Veröffentlicht: (2026)
Inference of Human-derived Specifications of Object Placement via Demonstration
von: Cuellar, Alex, et al.
Veröffentlicht: (2025)
von: Cuellar, Alex, et al.
Veröffentlicht: (2025)
SANGO: Socially Aware Navigation through Grouped Obstacles
von: Malladi, Rahath, et al.
Veröffentlicht: (2024)
von: Malladi, Rahath, et al.
Veröffentlicht: (2024)
$π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs
von: Wang, Siting, et al.
Veröffentlicht: (2026)
von: Wang, Siting, et al.
Veröffentlicht: (2026)
Learning Affordances at Inference-Time for Vision-Language-Action Models
von: Shah, Ameesh, et al.
Veröffentlicht: (2025)
von: Shah, Ameesh, et al.
Veröffentlicht: (2025)
Robotic 3D Flower Pose Estimation for Small-Scale Urban Farms
von: Muriki, Harsh, et al.
Veröffentlicht: (2025)
von: Muriki, Harsh, et al.
Veröffentlicht: (2025)
Towards Real-Time Interpolation for Enhanced AUV Deep Sea Mapping
von: Saxena, Devanshu
Veröffentlicht: (2025)
von: Saxena, Devanshu
Veröffentlicht: (2025)
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
von: Pai, Jonas, et al.
Veröffentlicht: (2025)
von: Pai, Jonas, et al.
Veröffentlicht: (2025)
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
von: Yuan, Haoran, et al.
Veröffentlicht: (2026)
von: Yuan, Haoran, et al.
Veröffentlicht: (2026)
CoordLight: Learning Decentralized Coordination for Network-Wide Traffic Signal Control
von: Zhang, Yifeng, et al.
Veröffentlicht: (2026)
von: Zhang, Yifeng, et al.
Veröffentlicht: (2026)
Slug Mobile: Test-Bench for RL Testing
von: Morris, Jonathan Wellington, et al.
Veröffentlicht: (2024)
von: Morris, Jonathan Wellington, et al.
Veröffentlicht: (2024)
Inference-Time Policy Steering through Human Interactions
von: Wang, Yanwei, et al.
Veröffentlicht: (2024)
von: Wang, Yanwei, et al.
Veröffentlicht: (2024)
Learning Contextually-Adaptive Rewards via Calibrated Features
von: Forsey-Smerek, Alexandra, et al.
Veröffentlicht: (2025)
von: Forsey-Smerek, Alexandra, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2025) -
FASTER: Rethinking Real-Time Flow VLAs
von: Lu, Yuxiang, et al.
Veröffentlicht: (2026) -
Running VLAs at Real-time Speed
von: Ma, Yunchao, et al.
Veröffentlicht: (2025) -
Realtime-VLA FLASH: Speculative Inference Framework for Diffusion-based VLAs
von: Niu, Jiahui, et al.
Veröffentlicht: (2026) -
VLAs are Confined yet Capable of Generalizing to Novel Instructions
von: Li, Quanyi
Veröffentlicht: (2025)