SAC-GLAM: Improving Online RL for LLM agents with Soft Actor-Critic and Hindsight Relabeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gaven, Loris, Romac, Clement, Carta, Thomas, Lamprier, Sylvain, Sigaud, Olivier, Oudeyer, Pierre-Yves |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents
von: Carta, Thomas, et al.
Veröffentlicht: (2025)
von: Carta, Thomas, et al.
Veröffentlicht: (2025)
MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces
von: Gaven, Loris, et al.
Veröffentlicht: (2025)
von: Gaven, Loris, et al.
Veröffentlicht: (2025)
Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning
von: Carta, Thomas, et al.
Veröffentlicht: (2023)
von: Carta, Thomas, et al.
Veröffentlicht: (2023)
Reinforcement Learning for Aligning Large Language Models Agents with Interactive Environments: Quantifying and Mitigating Prompt Overfitting
von: Aissi, Mohamed Salim, et al.
Veröffentlicht: (2024)
von: Aissi, Mohamed Salim, et al.
Veröffentlicht: (2024)
WorldLLM: Improving LLMs' world modeling using curiosity-driven theory-making
von: Levy, Guillaume, et al.
Veröffentlicht: (2025)
von: Levy, Guillaume, et al.
Veröffentlicht: (2025)
Imagine Beyond! Distributionally Robust Auto-Encoding for State Space Coverage in Online Reinforcement Learning
von: Castanet, Nicolas, et al.
Veröffentlicht: (2025)
von: Castanet, Nicolas, et al.
Veröffentlicht: (2025)
Autotelic Agents with Intrinsically Motivated Goal-Conditioned Reinforcement Learning: a Short Survey
von: Colas, Cédric, et al.
Veröffentlicht: (2020)
von: Colas, Cédric, et al.
Veröffentlicht: (2020)
CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement Learning
von: Colas, Cédric, et al.
Veröffentlicht: (2018)
von: Colas, Cédric, et al.
Veröffentlicht: (2018)
AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling
von: Ding, Liang
Veröffentlicht: (2026)
von: Ding, Liang
Veröffentlicht: (2026)
Improved Performances and Motivation in Intelligent Tutoring Systems: Combining Machine Learning and Learner Choice
von: Clément, Benjamin, et al.
Veröffentlicht: (2024)
von: Clément, Benjamin, et al.
Veröffentlicht: (2024)
Improving Zero-Shot Offline RL via Behavioral Task Sampling
von: Bendib, Nazim, et al.
Veröffentlicht: (2026)
von: Bendib, Nazim, et al.
Veröffentlicht: (2026)
Automatic Generation of Question Hints for Mathematics Problems using Large Language Models in Educational Technology
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2024)
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2024)
LogProber: Disentangling confidence from contamination in LLM responses
von: Yax, Nicolas, et al.
Veröffentlicht: (2024)
von: Yax, Nicolas, et al.
Veröffentlicht: (2024)
Self-Improving Language Models for Evolutionary Program Synthesis: A Case Study on ARC-AGI
von: Pourcel, Julien, et al.
Veröffentlicht: (2025)
von: Pourcel, Julien, et al.
Veröffentlicht: (2025)
PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling
von: Singh, Utsav, et al.
Veröffentlicht: (2024)
von: Singh, Utsav, et al.
Veröffentlicht: (2024)
PRISM: Perception Reasoning Interleaved for Sequential Decision Making
von: Aissi, Mohamed Salim, et al.
Veröffentlicht: (2026)
von: Aissi, Mohamed Salim, et al.
Veröffentlicht: (2026)
Optimizing for Persuasion Improves LLM Generalization: Evidence from Quality-Diversity Evolution of Debate Strategies
von: Reedi, Aksel Joonas, et al.
Veröffentlicht: (2025)
von: Reedi, Aksel Joonas, et al.
Veröffentlicht: (2025)
Reinforcement Learning Position Control of a Quadrotor Using Soft Actor-Critic (SAC)
von: Mahran, Youssef, et al.
Veröffentlicht: (2025)
von: Mahran, Youssef, et al.
Veröffentlicht: (2025)
DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty
von: Cui, Mingxuan, et al.
Veröffentlicht: (2025)
von: Cui, Mingxuan, et al.
Veröffentlicht: (2025)
A Definition of Open-Ended Learning Problems for Goal-Conditioned Agents
von: Sigaud, Olivier, et al.
Veröffentlicht: (2023)
von: Sigaud, Olivier, et al.
Veröffentlicht: (2023)
Exploring Flow-Lenia Universes with a Curiosity-driven AI Scientist: Discovering Diverse Ecosystem Dynamics
von: Michel, Thomas, et al.
Veröffentlicht: (2025)
von: Michel, Thomas, et al.
Veröffentlicht: (2025)
Stein Variational Black-Box Combinatorial Optimization
von: Landais, Thomas, et al.
Veröffentlicht: (2026)
von: Landais, Thomas, et al.
Veröffentlicht: (2026)
PhyloLM : Inferring the Phylogeny of Large Language Models and Predicting their Performances in Benchmarks
von: Yax, Nicolas, et al.
Veröffentlicht: (2024)
von: Yax, Nicolas, et al.
Veröffentlicht: (2024)
Collective Innovation in Groups of Large Language Models
von: Nisioti, Eleni, et al.
Veröffentlicht: (2024)
von: Nisioti, Eleni, et al.
Veröffentlicht: (2024)
Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment
von: Petitbois, Mathieu, et al.
Veröffentlicht: (2026)
von: Petitbois, Mathieu, et al.
Veröffentlicht: (2026)
Training Table Question Answering via SQL Query Decomposition
von: Mouravieff, Raphaël, et al.
Veröffentlicht: (2024)
von: Mouravieff, Raphaël, et al.
Veröffentlicht: (2024)
Structural Deep Encoding for Table Question Answering
von: Mouravieff, Raphaël, et al.
Veröffentlicht: (2025)
von: Mouravieff, Raphaël, et al.
Veröffentlicht: (2025)
Lyapunov Constrained Soft Actor-Critic (LC-SAC) using Koopman Operator Theory for Quadrotor Trajectory Tracking
von: Kushwaha, Dhruv S., et al.
Veröffentlicht: (2026)
von: Kushwaha, Dhruv S., et al.
Veröffentlicht: (2026)
Revisiting Discrete Soft Actor-Critic
von: Zhou, Haibin, et al.
Veröffentlicht: (2022)
von: Zhou, Haibin, et al.
Veröffentlicht: (2022)
Safe Langevin Soft Actor Critic
von: Keswani, Mahesh, et al.
Veröffentlicht: (2026)
von: Keswani, Mahesh, et al.
Veröffentlicht: (2026)
Average-Reward Soft Actor-Critic
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
Wasserstein Barycenter Soft Actor-Critic
von: Shahrooei, Zahra, et al.
Veröffentlicht: (2025)
von: Shahrooei, Zahra, et al.
Veröffentlicht: (2025)
Black-Box Combinatorial Optimization with Order-Invariant Reinforcement Learning
von: Goudet, Olivier, et al.
Veröffentlicht: (2025)
von: Goudet, Olivier, et al.
Veröffentlicht: (2025)
Agentic Adversarial QA for Improving Domain-Specific LLMs
von: Grari, Vincent, et al.
Veröffentlicht: (2026)
von: Grari, Vincent, et al.
Veröffentlicht: (2026)
FootstepNet: an Efficient Actor-Critic Method for Fast On-line Bipedal Footstep Planning and Forecasting
von: Gaspard, Clément, et al.
Veröffentlicht: (2024)
von: Gaspard, Clément, et al.
Veröffentlicht: (2024)
SAC-NeRF: Adaptive Ray Sampling for Neural Radiance Fields via Soft Actor-Critic Reinforcement Learning
von: Ge, Chenyu
Veröffentlicht: (2025)
von: Ge, Chenyu
Veröffentlicht: (2025)
Nature and Nature's God: A Philosophical and Scientific Defense of Aquinas's Unmoved Mover Argument. By DanielShields. Washington, D.C.: Catholic University of America Press, 2023. Pp. 328. $75.00.
von: Gaven Kerr
Veröffentlicht: (2024)
von: Gaven Kerr
Veröffentlicht: (2024)
Distributional Soft Actor-Critic with Three Refinements
von: Duan, Jingliang, et al.
Veröffentlicht: (2023)
von: Duan, Jingliang, et al.
Veröffentlicht: (2023)
Generative Actor-Critic with Soft Bridge Policies
von: He, Ke, et al.
Veröffentlicht: (2026)
von: He, Ke, et al.
Veröffentlicht: (2026)
Distributional Soft Actor-Critic with Diffusion Policy
von: Liu, Tong, et al.
Veröffentlicht: (2025)
von: Liu, Tong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents
von: Carta, Thomas, et al.
Veröffentlicht: (2025) -
MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces
von: Gaven, Loris, et al.
Veröffentlicht: (2025) -
Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning
von: Carta, Thomas, et al.
Veröffentlicht: (2023) -
Reinforcement Learning for Aligning Large Language Models Agents with Interactive Environments: Quantifying and Mitigating Prompt Overfitting
von: Aissi, Mohamed Salim, et al.
Veröffentlicht: (2024) -
WorldLLM: Improving LLMs' world modeling using curiosity-driven theory-making
von: Levy, Guillaume, et al.
Veröffentlicht: (2025)