Gespeichert in:
| Hauptverfasser: | Jiang, Daniel R., Bhandari, Jalaj, Yang, Yukai, Munos, Rémi, Lu, Tyler |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2511.21638 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Aligned Multi Objective Optimization
von: Efroni, Yonathan, et al.
Veröffentlicht: (2025)
von: Efroni, Yonathan, et al.
Veröffentlicht: (2025)
Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
von: Li, Junbo, et al.
Veröffentlicht: (2025)
von: Li, Junbo, et al.
Veröffentlicht: (2025)
Outcome-based Exploration for LLM Reasoning
von: Song, Yuda, et al.
Veröffentlicht: (2025)
von: Song, Yuda, et al.
Veröffentlicht: (2025)
On a few pitfalls in KL divergence gradient estimation for RL
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
Super-Exponential Regret for UCT, AlphaGo and Variants
von: Orseau, Laurent, et al.
Veröffentlicht: (2024)
von: Orseau, Laurent, et al.
Veröffentlicht: (2024)
Efficient RL Training for LLMs with Experience Replay
von: Arnal, Charles, et al.
Veröffentlicht: (2026)
von: Arnal, Charles, et al.
Veröffentlicht: (2026)
Bandits attack function optimization
von: Preux, Philippe, et al.
Veröffentlicht: (2026)
von: Preux, Philippe, et al.
Veröffentlicht: (2026)
Stochastic simultaneous optimistic optimization
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
RL-finetuning LLMs from on- and off-policy data with a single algorithm
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
Black-box optimization of noisy functions with unknown smoothness
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
Spectral Thompson sampling
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
VA-learning as a more efficient alternative to Q-learning
von: Tang, Yunhao, et al.
Veröffentlicht: (2023)
von: Tang, Yunhao, et al.
Veröffentlicht: (2023)
Spectral bandits for smooth graph functions
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Efficient learning by implicit exploration in bandit problems with side observations
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
Spectral bandits for smooth graph functions with applications in recommender systems
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)
Enhancing PPO with Trajectory-Aware Hybrid Policies
von: Liu, Qisai, et al.
Veröffentlicht: (2025)
von: Liu, Qisai, et al.
Veröffentlicht: (2025)
Mitigating Conversational Inertia in Multi-Turn Agents
von: Wan, Yang, et al.
Veröffentlicht: (2026)
von: Wan, Yang, et al.
Veröffentlicht: (2026)
Spectral bandits
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
von: Zhang, Qingru, et al.
Veröffentlicht: (2025)
von: Zhang, Qingru, et al.
Veröffentlicht: (2025)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
Eliciting Behaviors in Multi-Turn Conversations
von: Huang, Jing, et al.
Veröffentlicht: (2025)
von: Huang, Jing, et al.
Veröffentlicht: (2025)
Planning in entropy-regularized Markov decision processes and games
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model
von: Rowland, Mark, et al.
Veröffentlicht: (2024)
von: Rowland, Mark, et al.
Veröffentlicht: (2024)
Building Math Agents with Multi-Turn Iterative Preference Learning
von: Xiong, Wei, et al.
Veröffentlicht: (2024)
von: Xiong, Wei, et al.
Veröffentlicht: (2024)
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach
von: Zhang, Xinnan, et al.
Veröffentlicht: (2025)
von: Zhang, Xinnan, et al.
Veröffentlicht: (2025)
Sampling Complexity of TD and PPO in RKHS
von: Zou, Lu, et al.
Veröffentlicht: (2025)
von: Zou, Lu, et al.
Veröffentlicht: (2025)
Fix Initial Codes and Iteratively Refine Textual Directions Toward Safe Multi-Turn Code Correction
von: Tanaka, Yuto, et al.
Veröffentlicht: (2026)
von: Tanaka, Yuto, et al.
Veröffentlicht: (2026)
Multi-Turn Reasoning LLMs for Task Offloading in Mobile Edge Computing
von: Yang, Ning, et al.
Veröffentlicht: (2026)
von: Yang, Ning, et al.
Veröffentlicht: (2026)
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
von: Arnal, Charles, et al.
Veröffentlicht: (2025)
von: Arnal, Charles, et al.
Veröffentlicht: (2025)
Temporal Difference Flows
von: Farebrother, Jesse, et al.
Veröffentlicht: (2025)
von: Farebrother, Jesse, et al.
Veröffentlicht: (2025)
Asking Forever: Universal Activations Behind Turn Amplification in Conversational LLMs
von: Coalson, Zachary, et al.
Veröffentlicht: (2026)
von: Coalson, Zachary, et al.
Veröffentlicht: (2026)
VinePPO: Refining Credit Assignment in RL Training of LLMs
von: Kazemnejad, Amirhossein, et al.
Veröffentlicht: (2024)
von: Kazemnejad, Amirhossein, et al.
Veröffentlicht: (2024)
Directional-Clamp PPO
von: Karpel, Gilad, et al.
Veröffentlicht: (2025)
von: Karpel, Gilad, et al.
Veröffentlicht: (2025)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation
von: Feldman, Shai, et al.
Veröffentlicht: (2026)
von: Feldman, Shai, et al.
Veröffentlicht: (2026)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Aligned Multi Objective Optimization
von: Efroni, Yonathan, et al.
Veröffentlicht: (2025) -
Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
von: Li, Junbo, et al.
Veröffentlicht: (2025) -
Outcome-based Exploration for LLM Reasoning
von: Song, Yuda, et al.
Veröffentlicht: (2025) -
On a few pitfalls in KL divergence gradient estimation for RL
von: Tang, Yunhao, et al.
Veröffentlicht: (2025) -
Super-Exponential Regret for UCT, AlphaGo and Variants
von: Orseau, Laurent, et al.
Veröffentlicht: (2024)