AssistanceZero: Scalably Solving Assistance Games
Fuente:
arXiv
Saved in:
| Main Authors: | Laidlaw, Cassidy, Bronstein, Eli, Guo, Timothy, Feng, Dylan, Berglund, Lukas, Svegliato, Justin, Russell, Stuart, Dragan, Anca |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
by: Laidlaw, Cassidy, et al.
Published: (2023)
by: Laidlaw, Cassidy, et al.
Published: (2023)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
by: Laidlaw, Cassidy, et al.
Published: (2024)
by: Laidlaw, Cassidy, et al.
Published: (2024)
Bridging RL Theory and Practice with the Effective Horizon
by: Laidlaw, Cassidy, et al.
Published: (2023)
by: Laidlaw, Cassidy, et al.
Published: (2023)
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
by: Feng, Dylan, et al.
Published: (2026)
by: Feng, Dylan, et al.
Published: (2026)
Active teacher selection for reward learning
by: Freedman, Rachel, et al.
Published: (2023)
by: Freedman, Rachel, et al.
Published: (2023)
Observation Interference in Partially Observable Assistance Games
by: Emmons, Scott, et al.
Published: (2024)
by: Emmons, Scott, et al.
Published: (2024)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
by: Siththaranjan, Anand, et al.
Published: (2023)
by: Siththaranjan, Anand, et al.
Published: (2023)
Cooperative Inverse Reinforcement Learning
by: Hadfield-Menell, Dylan, et al.
Published: (2016)
by: Hadfield-Menell, Dylan, et al.
Published: (2016)
AI Alignment with Changing and Influenceable Reward Functions
by: Carroll, Micah, et al.
Published: (2024)
by: Carroll, Micah, et al.
Published: (2024)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
by: Lang, Leon, et al.
Published: (2024)
by: Lang, Leon, et al.
Published: (2024)
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
by: Ye, Yaowen, et al.
Published: (2025)
by: Ye, Yaowen, et al.
Published: (2025)
A Generalized Acquisition Function for Preference-based Reward Learning
by: Ellis, Evan, et al.
Published: (2024)
by: Ellis, Evan, et al.
Published: (2024)
Solving a Research Problem in Mathematical Statistics with AI Assistance
by: Dobriban, Edgar
Published: (2025)
by: Dobriban, Edgar
Published: (2025)
LLM Assistance for Pediatric Depression
by: Ignashina, Mariia, et al.
Published: (2025)
by: Ignashina, Mariia, et al.
Published: (2025)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
Adversaries Can Misuse Combinations of Safe Models
by: Jones, Erik, et al.
Published: (2024)
by: Jones, Erik, et al.
Published: (2024)
Deciphering AutoML Ensembles: cattleia's Assistance in Decision-Making
by: Kozak, Anna, et al.
Published: (2024)
by: Kozak, Anna, et al.
Published: (2024)
When to ASK: Uncertainty-Gated Language Assistance for Reinforcement Learning
by: Monteiro, Juarez, et al.
Published: (2026)
by: Monteiro, Juarez, et al.
Published: (2026)
On the Strengths and Weaknesses of Data for Open-set Embodied Assistance
by: Tambwekar, Pradyumna, et al.
Published: (2026)
by: Tambwekar, Pradyumna, et al.
Published: (2026)
Drift-aware Collaborative Assistance Mixture of Experts for Heterogeneous Multistream Learning
by: Yu, En, et al.
Published: (2025)
by: Yu, En, et al.
Published: (2025)
Shared Autonomy with IDA: Interventional Diffusion Assistance
by: McMahan, Brandon J., et al.
Published: (2024)
by: McMahan, Brandon J., et al.
Published: (2024)
Learning to Decide with AI Assistance under Human-Alignment
by: Benz, Nina Corvelo, et al.
Published: (2026)
by: Benz, Nina Corvelo, et al.
Published: (2026)
ColorGrid: A Multi-Agent Non-Stationary Environment for Goal Inference and Assistance
by: Risukhin, Andrey, et al.
Published: (2025)
by: Risukhin, Andrey, et al.
Published: (2025)
GDFlow: Anomaly Detection with NCDE-based Normalizing Flow for Advanced Driver Assistance System
by: Lee, Kangjun, et al.
Published: (2024)
by: Lee, Kangjun, et al.
Published: (2024)
Leveraging Knowledge Graphs and LLM Reasoning to Identify Operational Bottlenecks for Warehouse Planning Assistance
by: Parekh, Rishi, et al.
Published: (2025)
by: Parekh, Rishi, et al.
Published: (2025)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making
by: Myers, Vivek, et al.
Published: (2024)
by: Myers, Vivek, et al.
Published: (2024)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
Hide and Seek in Noise Labels: Noise-Robust Collaborative Active Learning with LLM-Powered Assistance
by: Yuan, Bo, et al.
Published: (2025)
by: Yuan, Bo, et al.
Published: (2025)
Value of Assistance for Grasping
by: Masarwy, Mohammad, et al.
Published: (2023)
by: Masarwy, Mohammad, et al.
Published: (2023)
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning
by: Li, Chengpeng, et al.
Published: (2024)
by: Li, Chengpeng, et al.
Published: (2024)
BanglaTalk: Towards Real-Time Speech Assistance for Bengali Regional Dialects
by: Hasan, Jakir, et al.
Published: (2025)
by: Hasan, Jakir, et al.
Published: (2025)
Training LLM Agents to Empower Humans
by: Ellis, Evan, et al.
Published: (2025)
by: Ellis, Evan, et al.
Published: (2025)
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
by: Williams, Marcus, et al.
Published: (2024)
by: Williams, Marcus, et al.
Published: (2024)
ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic Assistance
by: Sun, Liwen, et al.
Published: (2024)
by: Sun, Liwen, et al.
Published: (2024)
Pragmatic Instruction Following and Goal Assistance via Cooperative Language-Guided Inverse Planning
by: Zhi-Xuan, Tan, et al.
Published: (2024)
by: Zhi-Xuan, Tan, et al.
Published: (2024)
Aligning Robot and Human Representations
by: Bobu, Andreea, et al.
Published: (2023)
by: Bobu, Andreea, et al.
Published: (2023)
Multi-Surrogate-Teacher Assistance for Representation Alignment in Fingerprint-based Indoor Localization
by: Nguyen, Son Minh, et al.
Published: (2024)
by: Nguyen, Son Minh, et al.
Published: (2024)
Coprocessor Actor Critic: A Model-Based Reinforcement Learning Approach For Adaptive Brain Stimulation
by: Pan, Michelle, et al.
Published: (2024)
by: Pan, Michelle, et al.
Published: (2024)
Similar Items
-
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
by: Laidlaw, Cassidy, et al.
Published: (2023) -
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
by: Laidlaw, Cassidy, et al.
Published: (2024) -
Bridging RL Theory and Practice with the Effective Horizon
by: Laidlaw, Cassidy, et al.
Published: (2023) -
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
by: Feng, Dylan, et al.
Published: (2026) -
Active teacher selection for reward learning
by: Freedman, Rachel, et al.
Published: (2023)