LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
Fuente:
arXiv
Saved in:
| Main Authors: | Schmied, Thomas, Bornschein, Jörg, Grau-Moya, Jordi, Wulfmeier, Markus, Pascanu, Razvan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Retrieval-Augmented Decision Transformer: External Memory for In-context RL
by: Schmied, Thomas, et al.
Published: (2024)
by: Schmied, Thomas, et al.
Published: (2024)
Fine-Tuned In-Context Learners for Efficient Adaptation
by: Bornschein, Jorg, et al.
Published: (2025)
by: Bornschein, Jorg, et al.
Published: (2025)
Kalman Filter for Online Classification of Non-Stationary Data
by: Titsias, Michalis K., et al.
Published: (2023)
by: Titsias, Michalis K., et al.
Published: (2023)
A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks
by: Schmied, Thomas, et al.
Published: (2024)
by: Schmied, Thomas, et al.
Published: (2024)
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Meta-learning how to Share Credit among Macro-Actions
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
Revisiting Adam for Streaming Reinforcement Learning
by: Gogianu, Florin, et al.
Published: (2026)
by: Gogianu, Florin, et al.
Published: (2026)
Parameter Efficient Fine-tuning via Explained Variance Adaptation
by: Paischer, Fabian, et al.
Published: (2024)
by: Paischer, Fabian, et al.
Published: (2024)
Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning
by: Geles, Ismail, et al.
Published: (2026)
by: Geles, Ismail, et al.
Published: (2026)
Activated LoRA: Fine-tuned LLMs for Intrinsics
by: Greenewald, Kristjan, et al.
Published: (2025)
by: Greenewald, Kristjan, et al.
Published: (2025)
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024)
by: Veličković, Petar, et al.
Published: (2024)
On the generalization of language models from in-context learning and finetuning: a controlled study
by: Lampinen, Andrew K., et al.
Published: (2025)
by: Lampinen, Andrew K., et al.
Published: (2025)
Partition Tree Weighting for Non-Stationary Stochastic Bandits
by: Veness, Joel, et al.
Published: (2025)
by: Veness, Joel, et al.
Published: (2025)
State Soup: In-Context Skill Learning, Retrieval and Mixing
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
DARA: Few-shot Budget Allocation in Online Advertising via In-Context Decision Making with RL-Finetuned LLMs
by: Song, Mingxuan, et al.
Published: (2026)
by: Song, Mingxuan, et al.
Published: (2026)
Mining Generalizable Activation Functions
by: Vitvitskyi, Alex, et al.
Published: (2026)
by: Vitvitskyi, Alex, et al.
Published: (2026)
Perplexity Cannot Always Tell Right from Wrong
by: Veličković, Petar, et al.
Published: (2026)
by: Veličković, Petar, et al.
Published: (2026)
Layerwise LQR for Geometry-Aware Optimization of Deep Networks
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
Navigating Potholes with Geometry-Aware Sharpness Minimization
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
Growing Q-Networks: Solving Continuous Control Tasks with Adaptive Control Resolution
by: Seyde, Tim, et al.
Published: (2024)
by: Seyde, Tim, et al.
Published: (2024)
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Normalization and effective learning rates in reinforcement learning
by: Lyle, Clare, et al.
Published: (2024)
by: Lyle, Clare, et al.
Published: (2024)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
by: Li, Manling, et al.
Published: (2024)
by: Li, Manling, et al.
Published: (2024)
How do language models learn facts? Dynamics, curricula and hallucinations
by: Zucchet, Nicolas, et al.
Published: (2025)
by: Zucchet, Nicolas, et al.
Published: (2025)
How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition
by: Dong, Guanting, et al.
Published: (2023)
by: Dong, Guanting, et al.
Published: (2023)
Position Paper: Rethinking Privacy in RL for Sequential Decision-making in the Age of LLMs
by: Fan, Flint Xiaofeng, et al.
Published: (2025)
by: Fan, Flint Xiaofeng, et al.
Published: (2025)
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
by: Yu, Tongzhou, et al.
Published: (2025)
by: Yu, Tongzhou, et al.
Published: (2025)
Maxwell's Demon at Work: Efficient Pruning by Leveraging Saturation of Neurons
by: Dufort-Labbé, Simon, et al.
Published: (2024)
by: Dufort-Labbé, Simon, et al.
Published: (2024)
Asynchronous Algorithmic Alignment with Cocycles
by: Dudzik, Andrew, et al.
Published: (2023)
by: Dudzik, Andrew, et al.
Published: (2023)
Explaining RL Decisions with Trajectories
by: Deshmukh, Shripad Vilasrao, et al.
Published: (2023)
by: Deshmukh, Shripad Vilasrao, et al.
Published: (2023)
ZorBA: Zeroth-order Federated Fine-tuning of LLMs with Heterogeneous Block Activation
by: Meng, Chuiyang, et al.
Published: (2026)
by: Meng, Chuiyang, et al.
Published: (2026)
Understanding Prompt Tuning and In-Context Learning via Meta-Learning
by: Genewein, Tim, et al.
Published: (2025)
by: Genewein, Tim, et al.
Published: (2025)
PIANIST: Learning Partially Observable World Models with LLMs for Multi-Agent Decision Making
by: Light, Jonathan, et al.
Published: (2024)
by: Light, Jonathan, et al.
Published: (2024)
Beyond Greedy Exits: Improved Early Exit Decisions for Risk Control and Reliability
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
Torque-Aware Momentum
by: Malviya, Pranshu, et al.
Published: (2024)
by: Malviya, Pranshu, et al.
Published: (2024)
Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset
by: Galashov, Alexandre, et al.
Published: (2024)
by: Galashov, Alexandre, et al.
Published: (2024)
The Two-Stage Decision-Sampling Hypothesis: Understanding the Emergence of Self-Reflection in RL-Trained LLMs
by: Zhao, Zibo, et al.
Published: (2026)
by: Zhao, Zibo, et al.
Published: (2026)
Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
by: Chen, Jack, et al.
Published: (2025)
by: Chen, Jack, et al.
Published: (2025)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
by: Xu, Yinggan, et al.
Published: (2026)
by: Xu, Yinggan, et al.
Published: (2026)
Horizon Reduction Makes RL Scalable
by: Park, Seohong, et al.
Published: (2025)
by: Park, Seohong, et al.
Published: (2025)
Similar Items
-
Retrieval-Augmented Decision Transformer: External Memory for In-context RL
by: Schmied, Thomas, et al.
Published: (2024) -
Fine-Tuned In-Context Learners for Efficient Adaptation
by: Bornschein, Jorg, et al.
Published: (2025) -
Kalman Filter for Online Classification of Non-Stationary Data
by: Titsias, Michalis K., et al.
Published: (2023) -
A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks
by: Schmied, Thomas, et al.
Published: (2024) -
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025)