Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Carta, Thomas, Romac, Clément, Wolf, Thomas, Lamprier, Sylvain, Sigaud, Olivier, Oudeyer, Pierre-Yves |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reinforcement Learning for Aligning Large Language Models Agents with Interactive Environments: Quantifying and Mitigating Prompt Overfitting
von: Aissi, Mohamed Salim, et al.
Veröffentlicht: (2024)
von: Aissi, Mohamed Salim, et al.
Veröffentlicht: (2024)
SAC-GLAM: Improving Online RL for LLM agents with Soft Actor-Critic and Hindsight Relabeling
von: Gaven, Loris, et al.
Veröffentlicht: (2024)
von: Gaven, Loris, et al.
Veröffentlicht: (2024)
HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents
von: Carta, Thomas, et al.
Veröffentlicht: (2025)
von: Carta, Thomas, et al.
Veröffentlicht: (2025)
MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces
von: Gaven, Loris, et al.
Veröffentlicht: (2025)
von: Gaven, Loris, et al.
Veröffentlicht: (2025)
Imagine Beyond! Distributionally Robust Auto-Encoding for State Space Coverage in Online Reinforcement Learning
von: Castanet, Nicolas, et al.
Veröffentlicht: (2025)
von: Castanet, Nicolas, et al.
Veröffentlicht: (2025)
WorldLLM: Improving LLMs' world modeling using curiosity-driven theory-making
von: Levy, Guillaume, et al.
Veröffentlicht: (2025)
von: Levy, Guillaume, et al.
Veröffentlicht: (2025)
Autotelic Agents with Intrinsically Motivated Goal-Conditioned Reinforcement Learning: a Short Survey
von: Colas, Cédric, et al.
Veröffentlicht: (2020)
von: Colas, Cédric, et al.
Veröffentlicht: (2020)
Black-Box Combinatorial Optimization with Order-Invariant Reinforcement Learning
von: Goudet, Olivier, et al.
Veröffentlicht: (2025)
von: Goudet, Olivier, et al.
Veröffentlicht: (2025)
PhyloLM : Inferring the Phylogeny of Large Language Models and Predicting their Performances in Benchmarks
von: Yax, Nicolas, et al.
Veröffentlicht: (2024)
von: Yax, Nicolas, et al.
Veröffentlicht: (2024)
Reward-Preserving Attacks For Robust Reinforcement Learning
von: Schott, Lucas, et al.
Veröffentlicht: (2026)
von: Schott, Lucas, et al.
Veröffentlicht: (2026)
Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment
von: Petitbois, Mathieu, et al.
Veröffentlicht: (2026)
von: Petitbois, Mathieu, et al.
Veröffentlicht: (2026)
CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement Learning
von: Colas, Cédric, et al.
Veröffentlicht: (2018)
von: Colas, Cédric, et al.
Veröffentlicht: (2018)
Self-Improving Language Models for Evolutionary Program Synthesis: A Case Study on ARC-AGI
von: Pourcel, Julien, et al.
Veröffentlicht: (2025)
von: Pourcel, Julien, et al.
Veröffentlicht: (2025)
Physics-Informed Model and Hybrid Planning for Efficient Dyna-Style Reinforcement Learning
von: Asri, Zakariae El, et al.
Veröffentlicht: (2024)
von: Asri, Zakariae El, et al.
Veröffentlicht: (2024)
Improved Performances and Motivation in Intelligent Tutoring Systems: Combining Machine Learning and Learner Choice
von: Clément, Benjamin, et al.
Veröffentlicht: (2024)
von: Clément, Benjamin, et al.
Veröffentlicht: (2024)
RT-HCP: Dealing with Inference Delays and Sample Efficiency to Learn Directly on Robotic Platforms
von: Asri, Zakariae El, et al.
Veröffentlicht: (2025)
von: Asri, Zakariae El, et al.
Veröffentlicht: (2025)
Deinterleaving of Discrete Renewal Process Mixtures with Application to Electronic Support Measures
von: Pinsolle, Jean, et al.
Veröffentlicht: (2024)
von: Pinsolle, Jean, et al.
Veröffentlicht: (2024)
LogProber: Disentangling confidence from contamination in LLM responses
von: Yax, Nicolas, et al.
Veröffentlicht: (2024)
von: Yax, Nicolas, et al.
Veröffentlicht: (2024)
Stick to your Role! Stability of Personal Values Expressed in Large Language Models
von: Kovač, Grgur, et al.
Veröffentlicht: (2024)
von: Kovač, Grgur, et al.
Veröffentlicht: (2024)
Offline Learning of Controllable Diverse Behaviors
von: Petitbois, Mathieu, et al.
Veröffentlicht: (2025)
von: Petitbois, Mathieu, et al.
Veröffentlicht: (2025)
A Transformer Model for Predicting Chemical Products from Generic SMARTS Templates with Data Augmentation
von: Ozer, Derin, et al.
Veröffentlicht: (2025)
von: Ozer, Derin, et al.
Veröffentlicht: (2025)
ACES: Generating Diverse Programming Puzzles with with Autotelic Generative Models
von: Pourcel, Julien, et al.
Veröffentlicht: (2023)
von: Pourcel, Julien, et al.
Veröffentlicht: (2023)
Online Curvature-Aware Replay: Leveraging $\mathbf{2^{nd}}$ Order Information for Online Continual Learning
von: Urettini, Edoardo, et al.
Veröffentlicht: (2025)
von: Urettini, Edoardo, et al.
Veröffentlicht: (2025)
A tale of two goals: leveraging sequentiality in multi-goal scenarios
von: Serris, Olivier, et al.
Veröffentlicht: (2025)
von: Serris, Olivier, et al.
Veröffentlicht: (2025)
Discovering Sensorimotor Agency in Cellular Automata using Diversity Search
von: Hamon, Gautier, et al.
Veröffentlicht: (2024)
von: Hamon, Gautier, et al.
Veröffentlicht: (2024)
Navigation with QPHIL: Quantizing Planner for Hierarchical Implicit Q-Learning
von: Canesse, Alexi, et al.
Veröffentlicht: (2024)
von: Canesse, Alexi, et al.
Veröffentlicht: (2024)
Robust Deep Reinforcement Learning Through Adversarial Attacks and Training : A Survey
von: Schott, Lucas, et al.
Veröffentlicht: (2024)
von: Schott, Lucas, et al.
Veröffentlicht: (2024)
Automatic Generation of Question Hints for Mathematics Problems using Large Language Models in Educational Technology
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2024)
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2024)
Agentic Adversarial QA for Improving Domain-Specific LLMs
von: Grari, Vincent, et al.
Veröffentlicht: (2026)
von: Grari, Vincent, et al.
Veröffentlicht: (2026)
Environment Design for Inverse Reinforcement Learning
von: Buening, Thomas Kleine, et al.
Veröffentlicht: (2022)
von: Buening, Thomas Kleine, et al.
Veröffentlicht: (2022)
Zero-shot Model-based Reinforcement Learning using Large Language Models
von: Benechehab, Abdelhakim, et al.
Veröffentlicht: (2024)
von: Benechehab, Abdelhakim, et al.
Veröffentlicht: (2024)
Grounding Large Language Models In Embodied Environment With Imperfect World Models
von: Liu, Haolan, et al.
Veröffentlicht: (2024)
von: Liu, Haolan, et al.
Veröffentlicht: (2024)
Online Episodic Convex Reinforcement Learning
von: Moreno, Bianca Marin, et al.
Veröffentlicht: (2025)
von: Moreno, Bianca Marin, et al.
Veröffentlicht: (2025)
CaT: Constraints as Terminations for Legged Locomotion Reinforcement Learning
von: Chane-Sane, Elliot, et al.
Veröffentlicht: (2024)
von: Chane-Sane, Elliot, et al.
Veröffentlicht: (2024)
ACT: Agentic Classification Tree
von: Grari, Vincent, et al.
Veröffentlicht: (2025)
von: Grari, Vincent, et al.
Veröffentlicht: (2025)
Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX
von: Radji, Waris, et al.
Veröffentlicht: (2025)
von: Radji, Waris, et al.
Veröffentlicht: (2025)
Using Petri Nets as an Integrated Constraint Mechanism for Reinforcement Learning Tasks
von: Sachweh, Timon, et al.
Veröffentlicht: (2024)
von: Sachweh, Timon, et al.
Veröffentlicht: (2024)
Leveraging Environment Interaction for Automated PDDL Translation and Planning with Large Language Models
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2024)
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2024)
Online Continual Learning for Time Series: a Natural Score-driven Approach
von: Urettini, Edoardo, et al.
Veröffentlicht: (2026)
von: Urettini, Edoardo, et al.
Veröffentlicht: (2026)
Collective Innovation in Groups of Large Language Models
von: Nisioti, Eleni, et al.
Veröffentlicht: (2024)
von: Nisioti, Eleni, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reinforcement Learning for Aligning Large Language Models Agents with Interactive Environments: Quantifying and Mitigating Prompt Overfitting
von: Aissi, Mohamed Salim, et al.
Veröffentlicht: (2024) -
SAC-GLAM: Improving Online RL for LLM agents with Soft Actor-Critic and Hindsight Relabeling
von: Gaven, Loris, et al.
Veröffentlicht: (2024) -
HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents
von: Carta, Thomas, et al.
Veröffentlicht: (2025) -
MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces
von: Gaven, Loris, et al.
Veröffentlicht: (2025) -
Imagine Beyond! Distributionally Robust Auto-Encoding for State Space Coverage in Online Reinforcement Learning
von: Castanet, Nicolas, et al.
Veröffentlicht: (2025)