TROLL: Trust Regions improve Reinforcement Learning for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Becker, Philipp, Freymuth, Niklas, Thilges, Serge, Otto, Fabian, Neumann, Gerhard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KalMamba: Towards Efficient Probabilistic State Space Models for RL under Uncertainty
by: Becker, Philipp, et al.
Published: (2024)
by: Becker, Philipp, et al.
Published: (2024)
Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement Learning
by: Li, Ge, et al.
Published: (2024)
by: Li, Ge, et al.
Published: (2024)
MaNGO - Adaptable Graph Network Simulators via Meta-Learning
by: Dahlinger, Philipp, et al.
Published: (2025)
by: Dahlinger, Philipp, et al.
Published: (2025)
Adaptive Swarm Mesh Refinement using Deep Reinforcement Learning with Local Rewards
by: Freymuth, Niklas, et al.
Published: (2024)
by: Freymuth, Niklas, et al.
Published: (2024)
Can Neural Networks Provide Latent Embeddings for Telemetry-Aware Greedy Routing?
by: Boltres, Andreas, et al.
Published: (2026)
by: Boltres, Andreas, et al.
Published: (2026)
Efficient Off-Policy Learning for High-Dimensional Action Spaces
by: Otto, Fabian, et al.
Published: (2024)
by: Otto, Fabian, et al.
Published: (2024)
Combining Reconstruction and Contrastive Methods for Multimodal Representations in RL
by: Becker, Philipp, et al.
Published: (2023)
by: Becker, Philipp, et al.
Published: (2023)
Improving Long-Range Interactions in Graph Neural Simulators via Hamiltonian Dynamics
by: Hoang, Tai, et al.
Published: (2025)
by: Hoang, Tai, et al.
Published: (2025)
Diffusion-Based Hierarchical Graph Neural Networks for Simulating Nonlinear Solid Mechanics
by: Würth, Tobias, et al.
Published: (2025)
by: Würth, Tobias, et al.
Published: (2025)
Iterative Sizing Field Prediction for Adaptive Mesh Generation From Expert Demonstrations
by: Freymuth, Niklas, et al.
Published: (2024)
by: Freymuth, Niklas, et al.
Published: (2024)
Context-aware Learned Mesh-based Simulation via Trajectory-Level Meta-Learning
by: Dahlinger, Philipp, et al.
Published: (2025)
by: Dahlinger, Philipp, et al.
Published: (2025)
Learning Sub-Second Routing Optimization in Computer Networks requires Packet-Level Dynamics
by: Boltres, Andreas, et al.
Published: (2024)
by: Boltres, Andreas, et al.
Published: (2024)
Towards Near-Real-Time Telemetry-Aware Routing with Neural Routing Algorithms
by: Boltres, Andreas, et al.
Published: (2026)
by: Boltres, Andreas, et al.
Published: (2026)
PointPatchRL -- Masked Reconstruction Improves Reinforcement Learning on Point Clouds
by: Gyenes, Balázs, et al.
Published: (2024)
by: Gyenes, Balázs, et al.
Published: (2024)
Physics-informed MeshGraphNets (PI-MGNs): Neural finite element solvers for non-stationary and nonlinear simulations on arbitrary meshes
by: Würth, Tobias, et al.
Published: (2024)
by: Würth, Tobias, et al.
Published: (2024)
Point Cloud Sequence Encoding for Material-conditioned Graph Network Simulators
by: Dahlinger, Philipp, et al.
Published: (2026)
by: Dahlinger, Philipp, et al.
Published: (2026)
Movement Primitive Diffusion: Learning Gentle Robotic Manipulation of Deformable Objects
by: Scheikl, Paul Maria, et al.
Published: (2023)
by: Scheikl, Paul Maria, et al.
Published: (2023)
AMBER: Adaptive Mesh Generation by Iterative Mesh Resolution Prediction
by: Freymuth, Niklas, et al.
Published: (2025)
by: Freymuth, Niklas, et al.
Published: (2025)
Adaptive World Models: Learning Behaviors by Latent Imagination Under Non-Stationarity
by: Gospodinov, Emiliyan, et al.
Published: (2024)
by: Gospodinov, Emiliyan, et al.
Published: (2024)
Trust Region-Based Bayesian Optimisation to Discover Diverse Solutions
by: Perera, Kokila Kasuni, et al.
Published: (2025)
by: Perera, Kokila Kasuni, et al.
Published: (2025)
Trust Region Constrained Measure Transport in Path Space for Stochastic Optimal Control and Inference
by: Blessing, Denis, et al.
Published: (2025)
by: Blessing, Denis, et al.
Published: (2025)
Acquiring Diverse Skills using Curriculum Reinforcement Learning with Mixture of Experts
by: Celik, Onur, et al.
Published: (2024)
by: Celik, Onur, et al.
Published: (2024)
Integrating Human Knowledge Through Action Masking in Reinforcement Learning for Operations Research
by: Stappert, Mirko, et al.
Published: (2025)
by: Stappert, Mirko, et al.
Published: (2025)
Feasibility-Driven Trust Region Bayesian Optimization
by: Ascia, Paolo, et al.
Published: (2025)
by: Ascia, Paolo, et al.
Published: (2025)
Geometry-aware RL for Manipulation of Varying Shapes and Deformable Objects
by: Hoang, Tai, et al.
Published: (2025)
by: Hoang, Tai, et al.
Published: (2025)
Diffusion Policies creating a Trust Region for Offline Reinforcement Learning
by: Chen, Tianyu, et al.
Published: (2024)
by: Chen, Tianyu, et al.
Published: (2024)
Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning
by: Wijesundara, Chulabhaya, et al.
Published: (2026)
by: Wijesundara, Chulabhaya, et al.
Published: (2026)
Trust-Region Noise Search for Black-Box Alignment of Diffusion and Flow Models
by: Schweiger, Niklas, et al.
Published: (2026)
by: Schweiger, Niklas, et al.
Published: (2026)
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
Large Vision-Language Models for Knowledge-Grounded Data Annotation of Memes
by: Deng, Shiling, et al.
Published: (2025)
by: Deng, Shiling, et al.
Published: (2025)
Efficient Reinforcement Learning with Large Language Model Priors
by: Yan, Xue, et al.
Published: (2024)
by: Yan, Xue, et al.
Published: (2024)
Teaching Large Language Models to Reason with Reinforcement Learning
by: Havrilla, Alex, et al.
Published: (2024)
by: Havrilla, Alex, et al.
Published: (2024)
MuTT: A Multimodal Trajectory Transformer for Robot Skills
by: Kienle, Claudius, et al.
Published: (2024)
by: Kienle, Claudius, et al.
Published: (2024)
TOP-ERL: Transformer-based Off-Policy Episodic Reinforcement Learning
by: Li, Ge, et al.
Published: (2024)
by: Li, Ge, et al.
Published: (2024)
Is In-Context Learning in Large Language Models Bayesian? A Martingale Perspective
by: Falck, Fabian, et al.
Published: (2024)
by: Falck, Fabian, et al.
Published: (2024)
DIME:Diffusion-Based Maximum Entropy Reinforcement Learning
by: Celik, Onur, et al.
Published: (2025)
by: Celik, Onur, et al.
Published: (2025)
SEAR: Sample Efficient Action Chunking Reinforcement Learning
by: Nagy, C. F. Maximilian, et al.
Published: (2026)
by: Nagy, C. F. Maximilian, et al.
Published: (2026)
Reinforced Model Predictive Control via Trust-Region Quasi-Newton Policy Optimization
by: Brandner, Dean, et al.
Published: (2024)
by: Brandner, Dean, et al.
Published: (2024)
Reinforcement Learning Finetunes Small Subnetworks in Large Language Models
by: Mukherjee, Sagnik, et al.
Published: (2025)
by: Mukherjee, Sagnik, et al.
Published: (2025)
Trust Region Continual Learning as an Implicit Meta-Learner
by: Wang, Zekun, et al.
Published: (2026)
by: Wang, Zekun, et al.
Published: (2026)
Similar Items
-
KalMamba: Towards Efficient Probabilistic State Space Models for RL under Uncertainty
by: Becker, Philipp, et al.
Published: (2024) -
Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement Learning
by: Li, Ge, et al.
Published: (2024) -
MaNGO - Adaptable Graph Network Simulators via Meta-Learning
by: Dahlinger, Philipp, et al.
Published: (2025) -
Adaptive Swarm Mesh Refinement using Deep Reinforcement Learning with Local Rewards
by: Freymuth, Niklas, et al.
Published: (2024) -
Can Neural Networks Provide Latent Embeddings for Telemetry-Aware Greedy Routing?
by: Boltres, Andreas, et al.
Published: (2026)