Beyond Binary Success: Sample-Efficient and Statistically Rigorous Robot Policy Comparison
Fuente:
arXiv
Saved in:
| Main Authors: | Snyder, David, Badithela, Apurva, Matni, Nikolai, Pappas, George, Majumdar, Anirudha, Itkina, Masha, Nishimura, Haruki |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Is Your Imitation Learning Policy Better than Mine? Policy Comparison with Near-Optimal Stopping
by: Snyder, David, et al.
Published: (2025)
by: Snyder, David, et al.
Published: (2025)
How Generalizable Is My Behavior Cloning Policy? A Statistical Approach to Trustworthy Performance Evaluation
by: Vincent, Joseph A., et al.
Published: (2024)
by: Vincent, Joseph A., et al.
Published: (2024)
Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators
by: Badithela, Apurva, et al.
Published: (2025)
by: Badithela, Apurva, et al.
Published: (2025)
STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation
by: Goli, Hossein, et al.
Published: (2025)
by: Goli, Hossein, et al.
Published: (2025)
Explore until Confident: Efficient Exploration for Embodied Question Answering
by: Ren, Allen Z., et al.
Published: (2024)
by: Ren, Allen Z., et al.
Published: (2024)
Impact of Different Failures on a Robot's Perceived Reliability
by: Violette, Andrew, et al.
Published: (2026)
by: Violette, Andrew, et al.
Published: (2026)
SAFE: Multitask Failure Detection for Vision-Language-Action Models
by: Gu, Qiao, et al.
Published: (2025)
by: Gu, Qiao, et al.
Published: (2025)
Guiding Data Collection via Factored Scaling Curves
by: Zha, Lihan, et al.
Published: (2025)
by: Zha, Lihan, et al.
Published: (2025)
CUPID: Curating Data your Robot Loves with Influence Functions
by: Agia, Christopher, et al.
Published: (2025)
by: Agia, Christopher, et al.
Published: (2025)
PlayWorld: Learning Robot World Models from Autonomous Play
by: Yin, Tenny, et al.
Published: (2026)
by: Yin, Tenny, et al.
Published: (2026)
Privacy-Preserving Map-Free Exploration for Confirming the Absence of a Radioactive Source
by: Lepowsky, Eric, et al.
Published: (2024)
by: Lepowsky, Eric, et al.
Published: (2024)
Video Generation Models in Robotics -- Applications, Research Challenges, Future Directions
by: Mei, Zhiting, et al.
Published: (2026)
by: Mei, Zhiting, et al.
Published: (2026)
Benchmarking Particle Filter Algorithms for Efficient Velodyne-Based Vehicle Localization
by: Blanco-Claraco, Jose Luis, et al.
Published: (2024)
by: Blanco-Claraco, Jose Luis, et al.
Published: (2024)
Can We Detect Failures Without Failure Data? Uncertainty-Aware Runtime Failure Detection for Imitation Learning Policies
by: Xu, Chen, et al.
Published: (2025)
by: Xu, Chen, et al.
Published: (2025)
Deceptive Risk Minimization: Out-of-Distribution Generalization by Deceiving Distribution Shift Detectors
by: Majumdar, Anirudha
Published: (2025)
by: Majumdar, Anirudha
Published: (2025)
LOPR: Latent Occupancy PRediction using Generative Models
by: Lange, Bernard, et al.
Published: (2022)
by: Lange, Bernard, et al.
Published: (2022)
Learning-Based Fault Detection for Legged Robots in Remote Dynamic Environments
by: Stewart-Height, Abriana, et al.
Published: (2026)
by: Stewart-Height, Abriana, et al.
Published: (2026)
Learning Robot Safety from Sparse Human Feedback using Conformal Prediction
by: Feldman, Aaron O., et al.
Published: (2025)
by: Feldman, Aaron O., et al.
Published: (2025)
Exploratory Driving Performance and Car-Following Modeling for Autonomous Shuttles Based on Field Data
by: Favero, Renan, et al.
Published: (2023)
by: Favero, Renan, et al.
Published: (2023)
QuIP: Experimental design for expensive simulators with many Qualitative factors via Integer Programming
by: Liu, Yen-Chun, et al.
Published: (2025)
by: Liu, Yen-Chun, et al.
Published: (2025)
Integrated Scenario-based Analysis: A data-driven approach to support automated driving systems development and safety evaluation
by: Ali, Gibran, et al.
Published: (2024)
by: Ali, Gibran, et al.
Published: (2024)
COSMOS: A Data-Driven Probabilistic Time Series simulator for Chemical Plumes across Spatial Scales
by: Nag, Arunava, et al.
Published: (2025)
by: Nag, Arunava, et al.
Published: (2025)
Autonomous on-Demand Shuttles for First Mile-Last Mile Connectivity: Design, Optimization, and Impact Assessment
by: Roy, Sudipta, et al.
Published: (2024)
by: Roy, Sudipta, et al.
Published: (2024)
Composite Safety Potential Field for Highway Driving Risk Assessment
by: Zuo, Dachuan, et al.
Published: (2025)
by: Zuo, Dachuan, et al.
Published: (2025)
Self-supervised Multi-future Occupancy Forecasting for Autonomous Driving
by: Lange, Bernard, et al.
Published: (2024)
by: Lange, Bernard, et al.
Published: (2024)
Prognostic Framework for Robotic Manipulators Operating Under Dynamic Task Severities
by: Mohanty, Ayush, et al.
Published: (2024)
by: Mohanty, Ayush, et al.
Published: (2024)
Safe Planning in Interactive Environments via Iterative Policy Updates and Adversarially Robust Conformal Prediction
by: Mirzaeedodangeh, Omid, et al.
Published: (2025)
by: Mirzaeedodangeh, Omid, et al.
Published: (2025)
Predictive Red Teaming: Breaking Policies Without Breaking Robots
by: Majumdar, Anirudha, et al.
Published: (2025)
by: Majumdar, Anirudha, et al.
Published: (2025)
SciFi-Benchmark: Leveraging Science Fiction To Improve Robot Behavior
by: Sermanet, Pierre, et al.
Published: (2025)
by: Sermanet, Pierre, et al.
Published: (2025)
SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps
by: Shorinwa, Ola, et al.
Published: (2025)
by: Shorinwa, Ola, et al.
Published: (2025)
Reactive Temporal Logic-based Planning and Control for Interactive Robotic Tasks
by: Nawaz, Farhad, et al.
Published: (2024)
by: Nawaz, Farhad, et al.
Published: (2024)
Automated Vehicles at Unsignalized Intersections: Safety and Efficiency Implications of Mixed Human and Automated Traffic
by: Rahmani, Saeed, et al.
Published: (2024)
by: Rahmani, Saeed, et al.
Published: (2024)
Continuously Optimizing Radar Placement with Model Predictive Path Integrals
by: Potter, Michael, et al.
Published: (2024)
by: Potter, Michael, et al.
Published: (2024)
Markov Regime-Switching Intelligent Driver Model for Interpretable Car-Following Behavior
by: Zhang, Chengyuan, et al.
Published: (2025)
by: Zhang, Chengyuan, et al.
Published: (2025)
When Context Is Not Enough: Modeling Unexplained Variability in Car-Following Behavior
by: Zhang, Chengyuan, et al.
Published: (2025)
by: Zhang, Chengyuan, et al.
Published: (2025)
Contextual Safety Reasoning and Grounding for Open-World Robots
by: Ravichandran, Zachary, et al.
Published: (2026)
by: Ravichandran, Zachary, et al.
Published: (2026)
Uncertainty-Aware Deployment of Pre-trained Language-Conditioned Imitation Learning Policies
by: Wu, Bo, et al.
Published: (2024)
by: Wu, Bo, et al.
Published: (2024)
Exact Consistency Tests for Gaussian Mixture Filters using Normalized Deviation Squared Statistics
by: Ahmed, Nisar, et al.
Published: (2023)
by: Ahmed, Nisar, et al.
Published: (2023)
Domain Randomization is Sample Efficient for Linear Quadratic Control
by: Fujinami, Tesshu, et al.
Published: (2025)
by: Fujinami, Tesshu, et al.
Published: (2025)
Risk-Calibrated Human-Robot Interaction via Set-Valued Intent Prediction
by: Lidard, Justin, et al.
Published: (2024)
by: Lidard, Justin, et al.
Published: (2024)
Similar Items
-
Is Your Imitation Learning Policy Better than Mine? Policy Comparison with Near-Optimal Stopping
by: Snyder, David, et al.
Published: (2025) -
How Generalizable Is My Behavior Cloning Policy? A Statistical Approach to Trustworthy Performance Evaluation
by: Vincent, Joseph A., et al.
Published: (2024) -
Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators
by: Badithela, Apurva, et al.
Published: (2025) -
STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation
by: Goli, Hossein, et al.
Published: (2025) -
Explore until Confident: Efficient Exploration for Embodied Question Answering
by: Ren, Allen Z., et al.
Published: (2024)