Provable Interactive Learning with Hindsight Instruction Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Misra, Dipendra, Pacchiano, Aldo, Schapire, Robert E. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
von: Misra, Dipendra, et al.
Veröffentlicht: (2026)
von: Misra, Dipendra, et al.
Veröffentlicht: (2026)
Policy Improvement using Language Feedback Models
von: Zhong, Victor, et al.
Veröffentlicht: (2024)
von: Zhong, Victor, et al.
Veröffentlicht: (2024)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
von: Hong, Joey, et al.
Veröffentlicht: (2024)
von: Hong, Joey, et al.
Veröffentlicht: (2024)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
von: Wu, Yuning, et al.
Veröffentlicht: (2026)
von: Wu, Yuning, et al.
Veröffentlicht: (2026)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
Aligning LLM Agents by Learning Latent Preference from User Edits
von: Gao, Ge, et al.
Veröffentlicht: (2024)
von: Gao, Ge, et al.
Veröffentlicht: (2024)
When Less is Enough: Efficient Inference via Collaborative Reasoning
von: Chen, Yilei, et al.
Veröffentlicht: (2026)
von: Chen, Yilei, et al.
Veröffentlicht: (2026)
Second Order Bounds for Contextual Bandits with Function Approximation
von: Pacchiano, Aldo
Veröffentlicht: (2024)
von: Pacchiano, Aldo
Veröffentlicht: (2024)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
von: Hu, Michael Y., et al.
Veröffentlicht: (2025)
von: Hu, Michael Y., et al.
Veröffentlicht: (2025)
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2026)
Learning Rate-Free Reinforcement Learning: A Case for Model Selection with Non-Stationary Objectives
von: Afshar, Aida, et al.
Veröffentlicht: (2024)
von: Afshar, Aida, et al.
Veröffentlicht: (2024)
Dataset Reset Policy Optimization for RLHF
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
von: Lu, Zhicong, et al.
Veröffentlicht: (2026)
von: Lu, Zhicong, et al.
Veröffentlicht: (2026)
Improved Training Mechanism for Reinforcement Learning via Online Model Selection
von: Afshar, Aida, et al.
Veröffentlicht: (2025)
von: Afshar, Aida, et al.
Veröffentlicht: (2025)
On the Hardness of Bandit Learning
von: Brukhim, Nataly, et al.
Veröffentlicht: (2025)
von: Brukhim, Nataly, et al.
Veröffentlicht: (2025)
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
von: Yeo, Woongyeng, et al.
Veröffentlicht: (2026)
von: Yeo, Woongyeng, et al.
Veröffentlicht: (2026)
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
State-free Reinforcement Learning
von: Chen, Mingyu, et al.
Veröffentlicht: (2024)
von: Chen, Mingyu, et al.
Veröffentlicht: (2024)
In-Context Learning for Pure Exploration
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects
von: Latimer, Chris, et al.
Veröffentlicht: (2025)
von: Latimer, Chris, et al.
Veröffentlicht: (2025)
Interactive Training: Feedback-Driven Neural Network Optimization
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
Reinforcement Learning with Backtracking Feedback
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
von: Wang, Xingyao, et al.
Veröffentlicht: (2023)
von: Wang, Xingyao, et al.
Veröffentlicht: (2023)
Towards Principled Representation Learning from Videos for Reinforcement Learning
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
Data-Driven Online Model Selection With Regret Guarantees
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2023)
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2023)
Learning Personalized Agents from Human Feedback
von: Liang, Kaiqu, et al.
Veröffentlicht: (2026)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2026)
Distributionally Robust Reinforcement Learning with Human Feedback
von: Mandal, Debmalya, et al.
Veröffentlicht: (2025)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2025)
Learning to Reason from Feedback at Test-Time
von: Li, Yanyang, et al.
Veröffentlicht: (2025)
von: Li, Yanyang, et al.
Veröffentlicht: (2025)
Is Behavior Cloning All You Need? Understanding Horizon in Imitation Learning
von: Foster, Dylan J., et al.
Veröffentlicht: (2024)
von: Foster, Dylan J., et al.
Veröffentlicht: (2024)
Provable Length Generalization in Sequence Prediction via Spectral Filtering
von: Marsden, Annie, et al.
Veröffentlicht: (2024)
von: Marsden, Annie, et al.
Veröffentlicht: (2024)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
Parameter Efficient Reinforcement Learning from Human Feedback
von: Sidahmed, Hakim, et al.
Veröffentlicht: (2024)
von: Sidahmed, Hakim, et al.
Veröffentlicht: (2024)
Joint Learning of Context and Feedback Embeddings in Spoken Dialogue
von: Qian, Livia, et al.
Veröffentlicht: (2024)
von: Qian, Livia, et al.
Veröffentlicht: (2024)
RLVF: Learning from Verbal Feedback without Overgeneralization
von: Stephan, Moritz, et al.
Veröffentlicht: (2024)
von: Stephan, Moritz, et al.
Veröffentlicht: (2024)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
von: Chen, Yanxi, et al.
Veröffentlicht: (2024)
von: Chen, Yanxi, et al.
Veröffentlicht: (2024)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
von: Misra, Dipendra, et al.
Veröffentlicht: (2026) -
Policy Improvement using Language Feedback Models
von: Zhong, Victor, et al.
Veröffentlicht: (2024) -
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
von: Hong, Joey, et al.
Veröffentlicht: (2024) -
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
von: Wu, Yuning, et al.
Veröffentlicht: (2026) -
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)