Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Joey, Lin, Jessica, Dragan, Anca, Levine, Sergey |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL
by: Hong, Joey, et al.
Published: (2025)
by: Hong, Joey, et al.
Published: (2025)
Provable Interactive Learning with Hindsight Instruction Feedback
by: Misra, Dipendra, et al.
Published: (2024)
by: Misra, Dipendra, et al.
Published: (2024)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
by: Hu, Michael Y., et al.
Published: (2025)
by: Hu, Michael Y., et al.
Published: (2025)
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
by: Lu, Zhicong, et al.
Published: (2026)
by: Lu, Zhicong, et al.
Published: (2026)
Zero-Overhead Introspection for Adaptive Test-Time Compute
by: Manvi, Rohin, et al.
Published: (2025)
by: Manvi, Rohin, et al.
Published: (2025)
Learning to Model the World with Language
by: Lin, Jessy, et al.
Published: (2023)
by: Lin, Jessy, et al.
Published: (2023)
Learning to Assist Humans without Inferring Rewards
by: Myers, Vivek, et al.
Published: (2024)
by: Myers, Vivek, et al.
Published: (2024)
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
by: Yeo, Woongyeng, et al.
Published: (2026)
by: Yeo, Woongyeng, et al.
Published: (2026)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
by: Wu, Mian, et al.
Published: (2025)
by: Wu, Mian, et al.
Published: (2025)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
by: Zhou, Yifei, et al.
Published: (2024)
by: Zhou, Yifei, et al.
Published: (2024)
Training LLM Agents to Empower Humans
by: Ellis, Evan, et al.
Published: (2025)
by: Ellis, Evan, et al.
Published: (2025)
Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making
by: Myers, Vivek, et al.
Published: (2024)
by: Myers, Vivek, et al.
Published: (2024)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
by: Liang, Kaiqu, et al.
Published: (2025)
by: Liang, Kaiqu, et al.
Published: (2025)
UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
by: Zhai, Yuexiang, et al.
Published: (2024)
by: Zhai, Yuexiang, et al.
Published: (2024)
Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects
by: Latimer, Chris, et al.
Published: (2025)
by: Latimer, Chris, et al.
Published: (2025)
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
by: Butt, Natasha, et al.
Published: (2024)
by: Butt, Natasha, et al.
Published: (2024)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
by: Wu, Yuning, et al.
Published: (2026)
by: Wu, Yuning, et al.
Published: (2026)
Reinforcement Learning for Personalized Dialogue Management
by: Hengst, Floris den, et al.
Published: (2019)
by: Hengst, Floris den, et al.
Published: (2019)
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
Affective Multimodal Agents with Proactive Knowledge Grounding for Emotionally Aligned Marketing Dialogue
by: Yu, Lin, et al.
Published: (2025)
by: Yu, Lin, et al.
Published: (2025)
DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling
by: Wang, Minzheng, et al.
Published: (2024)
by: Wang, Minzheng, et al.
Published: (2024)
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
by: Guo, Ruohao, et al.
Published: (2025)
by: Guo, Ruohao, et al.
Published: (2025)
ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
by: Liu, Zexi, et al.
Published: (2025)
by: Liu, Zexi, et al.
Published: (2025)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
Dialogue Ontology Relation Extraction via Constrained Chain-of-Thought Decoding
by: Vukovic, Renato, et al.
Published: (2024)
by: Vukovic, Renato, et al.
Published: (2024)
HIPO: Instruction Hierarchy via Constrained Reinforcement Learning
by: Chen, Keru, et al.
Published: (2026)
by: Chen, Keru, et al.
Published: (2026)
MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning
by: Chen, Guoxin, et al.
Published: (2025)
by: Chen, Guoxin, et al.
Published: (2025)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
by: Lieberum, Tom, et al.
Published: (2024)
by: Lieberum, Tom, et al.
Published: (2024)
Unfamiliar Finetuning Examples Control How Language Models Hallucinate
by: Kang, Katie, et al.
Published: (2024)
by: Kang, Katie, et al.
Published: (2024)
Coprocessor Actor Critic: A Model-Based Reinforcement Learning Approach For Adaptive Brain Stimulation
by: Pan, Michelle, et al.
Published: (2024)
by: Pan, Michelle, et al.
Published: (2024)
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
by: Guo, Yiju, et al.
Published: (2026)
by: Guo, Yiju, et al.
Published: (2026)
Learning Retrieval Augmentation for Personalized Dialogue Generation
by: Huang, Qiushi, et al.
Published: (2024)
by: Huang, Qiushi, et al.
Published: (2024)
CLAIM: An Intent-Driven Multi-Agent Framework for Analyzing Manipulation in Courtroom Dialogues
by: Sheshanarayana, Disha, et al.
Published: (2025)
by: Sheshanarayana, Disha, et al.
Published: (2025)
Adversarial Reinforcement Learning for Large Language Model Agent Safety
by: Wang, Zizhao, et al.
Published: (2025)
by: Wang, Zizhao, et al.
Published: (2025)
Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions
by: Tan, Cheng, et al.
Published: (2024)
by: Tan, Cheng, et al.
Published: (2024)
Joint Learning of Context and Feedback Embeddings in Spoken Dialogue
by: Qian, Livia, et al.
Published: (2024)
by: Qian, Livia, et al.
Published: (2024)
Similar Items
-
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
by: Hong, Joey, et al.
Published: (2024) -
Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL
by: Hong, Joey, et al.
Published: (2025) -
Provable Interactive Learning with Hindsight Instruction Feedback
by: Misra, Dipendra, et al.
Published: (2024) -
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
by: Hu, Michael Y., et al.
Published: (2025) -
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
by: Lu, Zhicong, et al.
Published: (2026)