Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Singhi, Nishad, Bialas, Christian, Jauhri, Snehal, Prasad, Vignesh, Chalvatzaki, Georgia, Rohrbach, Marcus, Rohrbach, Anna |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
by: Singhi, Nishad, et al.
Published: (2025)
by: Singhi, Nishad, et al.
Published: (2025)
Active-Perceptive Motion Generation for Mobile Manipulation
by: Jauhri, Snehal, et al.
Published: (2023)
by: Jauhri, Snehal, et al.
Published: (2023)
Whole-Body Mobile Manipulation using Offline Reinforcement Learning on Sub-optimal Controllers
by: Jauhri, Snehal, et al.
Published: (2026)
by: Jauhri, Snehal, et al.
Published: (2026)
2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human Videos
by: Heidinger, Marvin, et al.
Published: (2025)
by: Heidinger, Marvin, et al.
Published: (2025)
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026)
by: Rothermel, Mark, et al.
Published: (2026)
Geometric Erasure by Contrastive Velocity Matching in Rectified Flows
by: Grebe, Jonas Henry, et al.
Published: (2026)
by: Grebe, Jonas Henry, et al.
Published: (2026)
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
by: Braun, Tobias, et al.
Published: (2025)
by: Braun, Tobias, et al.
Published: (2025)
SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring
by: Rodriguez, Hector G., et al.
Published: (2026)
by: Rodriguez, Hector G., et al.
Published: (2026)
Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models
by: Braun, Tobias, et al.
Published: (2026)
by: Braun, Tobias, et al.
Published: (2026)
6DOPE-GS: Online 6D Object Pose Estimation using Gaussian Splatting
by: Jin, Yufeng, et al.
Published: (2024)
by: Jin, Yufeng, et al.
Published: (2024)
Learning Any-View 6DoF Robotic Grasping in Cluttered Scenes via Neural Surface Rendering
by: Jauhri, Snehal, et al.
Published: (2023)
by: Jauhri, Snehal, et al.
Published: (2023)
ActionFlow: Equivariant, Accurate, and Efficient Policies with Spatially Symmetric Flow Matching
by: Funk, Niklas, et al.
Published: (2024)
by: Funk, Niklas, et al.
Published: (2024)
EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection
by: Marik, Aritra, et al.
Published: (2026)
by: Marik, Aritra, et al.
Published: (2026)
UniFField: A Generalizable Unified Neural Feature Field for Visual, Semantic, and Spatial Uncertainties in Any Scene
by: Maurer, Christian, et al.
Published: (2025)
by: Maurer, Christian, et al.
Published: (2025)
Think Twice, Act Once: A Co-Evolution Framework of LLM and RL for Large-Scale Decision Making
by: Wan, Xu, et al.
Published: (2025)
by: Wan, Xu, et al.
Published: (2025)
DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection
by: Klemt, Marcel, et al.
Published: (2025)
by: Klemt, Marcel, et al.
Published: (2025)
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
by: Wieczorek, Tobias Jan, et al.
Published: (2025)
by: Wieczorek, Tobias Jan, et al.
Published: (2025)
Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection
by: Phan, Hoang, et al.
Published: (2025)
by: Phan, Hoang, et al.
Published: (2025)
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
by: Braun, Tobias, et al.
Published: (2024)
by: Braun, Tobias, et al.
Published: (2024)
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
by: Abdessaied, Adnen, et al.
Published: (2025)
by: Abdessaied, Adnen, et al.
Published: (2025)
Bimanual Robot Manipulation via Multi-Agent In-Context Learning
by: Palma, Alessio, et al.
Published: (2026)
by: Palma, Alessio, et al.
Published: (2026)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
An Adaptive Multi Agent Bitcoin Trading System
by: Singhi, Aadi
Published: (2025)
by: Singhi, Aadi
Published: (2025)
When Do Diffusion Models learn to Generate Multiple Objects?
by: Jeong, Yujin, et al.
Published: (2026)
by: Jeong, Yujin, et al.
Published: (2026)
Chrono: A Simple Blueprint for Representing Time in MLLMs
by: Rodriguez, Hector, et al.
Published: (2024)
by: Rodriguez, Hector, et al.
Published: (2024)
Efficient Pre-training for Localized Instruction Generation of Videos
by: Batra, Anil, et al.
Published: (2023)
by: Batra, Anil, et al.
Published: (2023)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
by: Singhi, Nishad, et al.
Published: (2024)
by: Singhi, Nishad, et al.
Published: (2024)
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
by: Kurz, Paul Jonas, et al.
Published: (2026)
by: Kurz, Paul Jonas, et al.
Published: (2026)
Shape-Guided Diffusion with Inside-Outside Attention
by: Park, Dong Huk, et al.
Published: (2022)
by: Park, Dong Huk, et al.
Published: (2022)
See and Think: Embodied Agent in Virtual Environment
by: Zhao, Zhonghan, et al.
Published: (2023)
by: Zhao, Zhonghan, et al.
Published: (2023)
En búsqueda de un cuidado universal y cultural
by: Cecilia Rohrbach
Published: (2007)
by: Cecilia Rohrbach
Published: (2007)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
by: Lacombe, Romain, et al.
Published: (2025)
by: Lacombe, Romain, et al.
Published: (2025)
ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement
by: Jiao, Difan, et al.
Published: (2026)
by: Jiao, Difan, et al.
Published: (2026)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
by: Zohrabi, Reihaneh, et al.
Published: (2026)
by: Zohrabi, Reihaneh, et al.
Published: (2026)
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
by: Yang, Rui, et al.
Published: (2026)
by: Yang, Rui, et al.
Published: (2026)
Uncertainty in Action: Confidence Elicitation in Embodied Agents
by: Yu, Tianjiao, et al.
Published: (2025)
by: Yu, Tianjiao, et al.
Published: (2025)
Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models
by: Tan, Xudong, et al.
Published: (2025)
by: Tan, Xudong, et al.
Published: (2025)
Parallax: Why AI Agents That Think Must Never Act
by: Fokou, Joel
Published: (2026)
by: Fokou, Joel
Published: (2026)
Think Twice: Measuring the Efficiency of Eliminating Prediction Shortcuts of Question Answering Models
by: Mikula, Lukáš, et al.
Published: (2023)
by: Mikula, Lukáš, et al.
Published: (2023)
Semantic-Geometric Task Representations for Bimanual Manipulation from Human Demonstrations to Robot Action Planning
by: Herbert, Franziska, et al.
Published: (2026)
by: Herbert, Franziska, et al.
Published: (2026)
Similar Items
-
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
by: Singhi, Nishad, et al.
Published: (2025) -
Active-Perceptive Motion Generation for Mobile Manipulation
by: Jauhri, Snehal, et al.
Published: (2023) -
Whole-Body Mobile Manipulation using Offline Reinforcement Learning on Sub-optimal Controllers
by: Jauhri, Snehal, et al.
Published: (2026) -
2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human Videos
by: Heidinger, Marvin, et al.
Published: (2025) -
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026)