Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Rachel, Qu, Jingyi, Bobu, Andreea, Hadfield-Menell, Dylan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Goal Inference from Open-Ended Dialog
by: Ma, Rachel, et al.
Published: (2024)
by: Ma, Rachel, et al.
Published: (2024)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
by: Hahm, Dongyoon, et al.
Published: (2026)
by: Hahm, Dongyoon, et al.
Published: (2026)
Robots That Know What to Ask: Recovering Misaligned Rewards through Targeted Explanations
by: Merker, Helena, et al.
Published: (2026)
by: Merker, Helena, et al.
Published: (2026)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
by: Ma, Rachel, et al.
Published: (2026)
by: Ma, Rachel, et al.
Published: (2026)
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
by: Damani, Mehul, et al.
Published: (2024)
by: Damani, Mehul, et al.
Published: (2024)
Aligning Robot and Human Representations
by: Bobu, Andreea, et al.
Published: (2023)
by: Bobu, Andreea, et al.
Published: (2023)
The Autonomy-Alignment Problem in Open-Ended Learning Robots: Formalising the Purpose Framework
by: Baldassarre, Gianluca, et al.
Published: (2024)
by: Baldassarre, Gianluca, et al.
Published: (2024)
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models
by: Sankaranarayanan, Aruna, et al.
Published: (2025)
by: Sankaranarayanan, Aruna, et al.
Published: (2025)
Enabling Adaptive Agent Training in Open-Ended Simulators by Targeting Diversity
by: Costales, Robby, et al.
Published: (2024)
by: Costales, Robby, et al.
Published: (2024)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
by: Siththaranjan, Anand, et al.
Published: (2023)
by: Siththaranjan, Anand, et al.
Published: (2023)
AutoLibra: Agent Metric Induction from Open-Ended Human Feedback
by: Zhu, Hao, et al.
Published: (2025)
by: Zhu, Hao, et al.
Published: (2025)
MaxMin-RLHF: Alignment with Diverse Human Preferences
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
by: Chen, Hui, et al.
Published: (2025)
by: Chen, Hui, et al.
Published: (2025)
Prompt Injection as Role Confusion
by: Ye, Charles, et al.
Published: (2026)
by: Ye, Charles, et al.
Published: (2026)
LEGENT: Open Platform for Embodied Agents
by: Cheng, Zhili, et al.
Published: (2024)
by: Cheng, Zhili, et al.
Published: (2024)
Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
by: Zheng, Qinqing, et al.
Published: (2024)
by: Zheng, Qinqing, et al.
Published: (2024)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026)
by: Lee, Jaehyeok, et al.
Published: (2026)
Mental Modeling of Reinforcement Learning Agents by Language Models
by: Lu, Wenhao, et al.
Published: (2024)
by: Lu, Wenhao, et al.
Published: (2024)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
by: Li, Manling, et al.
Published: (2024)
by: Li, Manling, et al.
Published: (2024)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
Dreaming in Code for Curriculum Learning in Open-Ended Worlds
by: Mitsides, Konstantinos, et al.
Published: (2026)
by: Mitsides, Konstantinos, et al.
Published: (2026)
MCU: An Evaluation Framework for Open-Ended Game Agents
by: Zheng, Xinyue, et al.
Published: (2023)
by: Zheng, Xinyue, et al.
Published: (2023)
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
by: Chen, Kaiyuan, et al.
Published: (2025)
by: Chen, Kaiyuan, et al.
Published: (2025)
DIVERGE: Diversity-Enhanced RAG for Open-Ended Information Seeking
by: Hu, Tianyi, et al.
Published: (2026)
by: Hu, Tianyi, et al.
Published: (2026)
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
by: Samvelyan, Mikayel, et al.
Published: (2024)
by: Samvelyan, Mikayel, et al.
Published: (2024)
An Empirical Study on Context Length for Open-Domain Dialog Generation
by: Shen, Xinyi, et al.
Published: (2024)
by: Shen, Xinyi, et al.
Published: (2024)
Preference-Conditioned Language-Guided Abstraction
by: Peng, Andi, et al.
Published: (2024)
by: Peng, Andi, et al.
Published: (2024)
GRLO: Towards Generalizable Reinforcement Learning in Open-Ended Environments from Zero
by: Yin, Shangjian, et al.
Published: (2026)
by: Yin, Shangjian, et al.
Published: (2026)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
by: Krishna, Kundan, et al.
Published: (2025)
by: Krishna, Kundan, et al.
Published: (2025)
HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks
by: Wang, Zhilin, et al.
Published: (2025)
by: Wang, Zhilin, et al.
Published: (2025)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
by: Ye, Zhiling, et al.
Published: (2025)
by: Ye, Zhiling, et al.
Published: (2025)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
by: Lu, Chris, et al.
Published: (2024)
by: Lu, Chris, et al.
Published: (2024)
FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making
by: Wang, Yucen, et al.
Published: (2025)
by: Wang, Yucen, et al.
Published: (2025)
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
GRAIL: Goal Recognition Alignment through Imitation Learning
by: Elhadad, Osher, et al.
Published: (2026)
by: Elhadad, Osher, et al.
Published: (2026)
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
by: Wang, Zhanliang, et al.
Published: (2026)
by: Wang, Zhanliang, et al.
Published: (2026)
QuickLAP: Quick Language-Action Preference Learning for Semi-Autonomous Agents
by: Nader, Jordan Abi, et al.
Published: (2025)
by: Nader, Jordan Abi, et al.
Published: (2025)
Human-Guided Harm Recovery for Computer Use Agents
by: Li, Christy, et al.
Published: (2026)
by: Li, Christy, et al.
Published: (2026)
The Essential Role of Causality in Foundation World Models for Embodied AI
by: Gupta, Tarun, et al.
Published: (2024)
by: Gupta, Tarun, et al.
Published: (2024)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
by: Casper, Stephen, et al.
Published: (2024)
by: Casper, Stephen, et al.
Published: (2024)
Similar Items
-
Goal Inference from Open-Ended Dialog
by: Ma, Rachel, et al.
Published: (2024) -
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
by: Hahm, Dongyoon, et al.
Published: (2026) -
Robots That Know What to Ask: Recovering Misaligned Rewards through Targeted Explanations
by: Merker, Helena, et al.
Published: (2026) -
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
by: Ma, Rachel, et al.
Published: (2026) -
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
by: Damani, Mehul, et al.
Published: (2024)