RL in Name Only? Analyzing the Structural Assumptions in RL post-training for LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Samineni, Soumya Rani, Kalwar, Durgesh, Valmeekam, Karthik, Stechly, Kaya, Kambhampati, Subbarao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
von: Samineni, Soumya Rani, et al.
Veröffentlicht: (2025)
von: Samineni, Soumya Rani, et al.
Veröffentlicht: (2025)
Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
von: Kambhampati, Subbarao, et al.
Veröffentlicht: (2025)
von: Kambhampati, Subbarao, et al.
Veröffentlicht: (2025)
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
von: Valmeekam, Karthik, et al.
Veröffentlicht: (2024)
von: Valmeekam, Karthik, et al.
Veröffentlicht: (2024)
Chain of Thoughtlessness? An Analysis of CoT in Planning
von: Stechly, Kaya, et al.
Veröffentlicht: (2024)
von: Stechly, Kaya, et al.
Veröffentlicht: (2024)
On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks
von: Stechly, Kaya, et al.
Veröffentlicht: (2024)
von: Stechly, Kaya, et al.
Veröffentlicht: (2024)
Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
von: Valmeekam, Karthik, et al.
Veröffentlicht: (2025)
von: Valmeekam, Karthik, et al.
Veröffentlicht: (2025)
Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
von: Palod, Vardhan, et al.
Veröffentlicht: (2025)
von: Palod, Vardhan, et al.
Veröffentlicht: (2025)
Planning in Strawberry Fields: Evaluating and Improving the Planning and Scheduling Capabilities of LRM o1
von: Valmeekam, Karthik, et al.
Veröffentlicht: (2024)
von: Valmeekam, Karthik, et al.
Veröffentlicht: (2024)
LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
von: Kambhampati, Subbarao, et al.
Veröffentlicht: (2024)
von: Kambhampati, Subbarao, et al.
Veröffentlicht: (2024)
(How) Do reasoning models reason?
von: Subbarao Kambhampati, et al.
Veröffentlicht: (2025)
von: Subbarao Kambhampati, et al.
Veröffentlicht: (2025)
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning
von: Bhambri, Siddhant, et al.
Veröffentlicht: (2024)
von: Bhambri, Siddhant, et al.
Veröffentlicht: (2024)
Can Large Language Models Reason and Plan?
von: Kambhampati, Subbarao
Veröffentlicht: (2024)
von: Kambhampati, Subbarao
Veröffentlicht: (2024)
Mind The Gap: Quantifying Mechanistic Gaps in Algorithmic Reasoning via Neural Compilation
von: Saldyt, Lucas, et al.
Veröffentlicht: (2025)
von: Saldyt, Lucas, et al.
Veröffentlicht: (2025)
Efficiency Boost in Decentralized Optimization: Reimagining Neighborhood Aggregation with Minimal Overhead
von: Kalwar, Durgesh, et al.
Veröffentlicht: (2025)
von: Kalwar, Durgesh, et al.
Veröffentlicht: (2025)
Robust Planning with Compound LLM Architectures: An LLM-Modulo Approach
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
von: Xiao, Wei, et al.
Veröffentlicht: (2025)
von: Xiao, Wei, et al.
Veröffentlicht: (2025)
Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
von: Zhang, Yiqi, et al.
Veröffentlicht: (2026)
von: Zhang, Yiqi, et al.
Veröffentlicht: (2026)
Partial Policy Gradients for RL in LLMs
von: Mathur, Puneet, et al.
Veröffentlicht: (2026)
von: Mathur, Puneet, et al.
Veröffentlicht: (2026)
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
von: Ran-Milo, Yuval, et al.
Veröffentlicht: (2026)
von: Ran-Milo, Yuval, et al.
Veröffentlicht: (2026)
A Systematic Investigation of The RL-Jailbreaker in LLMs
von: Mohammedalamen, Montaser, et al.
Veröffentlicht: (2026)
von: Mohammedalamen, Montaser, et al.
Veröffentlicht: (2026)
Prioritized Replay for RL Post-training
von: Fatemi, Mehdi
Veröffentlicht: (2026)
von: Fatemi, Mehdi
Veröffentlicht: (2026)
RL$^3$: Boosting Meta Reinforcement Learning via RL inside RL$^2$
von: Bhatia, Abhinav, et al.
Veröffentlicht: (2023)
von: Bhatia, Abhinav, et al.
Veröffentlicht: (2023)
RL-Struct: A Lightweight Reinforcement Learning Framework for Reliable Structured Output in LLMs
von: Hu, Ruike, et al.
Veröffentlicht: (2025)
von: Hu, Ruike, et al.
Veröffentlicht: (2025)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
von: Mark, Max Sobol, et al.
Veröffentlicht: (2024)
von: Mark, Max Sobol, et al.
Veröffentlicht: (2024)
Toward Explainable Offline RL: Analyzing Representations in Intrinsically Motivated Decision Transformers
von: Guiducci, Leonardo, et al.
Veröffentlicht: (2025)
von: Guiducci, Leonardo, et al.
Veröffentlicht: (2025)
Incorporating Human Flexibility through Reward Preferences in Human-AI Teaming
von: Bhambri, Siddhant, et al.
Veröffentlicht: (2023)
von: Bhambri, Siddhant, et al.
Veröffentlicht: (2023)
A Comedy of Estimators: On KL Regularization in RL Training of LLMs
von: Shah, Vedant, et al.
Veröffentlicht: (2025)
von: Shah, Vedant, et al.
Veröffentlicht: (2025)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
von: Wu, Runzhe, et al.
Veröffentlicht: (2025)
von: Wu, Runzhe, et al.
Veröffentlicht: (2025)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
von: Pignatelli, Eduardo, et al.
Veröffentlicht: (2024)
von: Pignatelli, Eduardo, et al.
Veröffentlicht: (2024)
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2026)
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2026)
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
von: Huang, Luke J., et al.
Veröffentlicht: (2026)
von: Huang, Luke J., et al.
Veröffentlicht: (2026)
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
von: Su, Jianhai, et al.
Veröffentlicht: (2025)
von: Su, Jianhai, et al.
Veröffentlicht: (2025)
Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward
von: Wen, Xuexiang, et al.
Veröffentlicht: (2026)
von: Wen, Xuexiang, et al.
Veröffentlicht: (2026)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
von: Chen, Zihan, et al.
Veröffentlicht: (2025)
von: Chen, Zihan, et al.
Veröffentlicht: (2025)
Combining LLM decision and RL action selection to improve RL policy for adaptive interventions
von: Karine, Karine, et al.
Veröffentlicht: (2025)
von: Karine, Karine, et al.
Veröffentlicht: (2025)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
von: Muslimani, Calarina, et al.
Veröffentlicht: (2025)
von: Muslimani, Calarina, et al.
Veröffentlicht: (2025)
RL-Finetuned LLMs for Privacy-Preserving Synthetic Rewriting
von: Shi, Zhan, et al.
Veröffentlicht: (2025)
von: Shi, Zhan, et al.
Veröffentlicht: (2025)
EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation
von: Shi, Jiahe, et al.
Veröffentlicht: (2025)
von: Shi, Jiahe, et al.
Veröffentlicht: (2025)
Position Paper: Rethinking Privacy in RL for Sequential Decision-making in the Age of LLMs
von: Fan, Flint Xiaofeng, et al.
Veröffentlicht: (2025)
von: Fan, Flint Xiaofeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
von: Samineni, Soumya Rani, et al.
Veröffentlicht: (2025) -
Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
von: Kambhampati, Subbarao, et al.
Veröffentlicht: (2025) -
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
von: Valmeekam, Karthik, et al.
Veröffentlicht: (2024) -
Chain of Thoughtlessness? An Analysis of CoT in Planning
von: Stechly, Kaya, et al.
Veröffentlicht: (2024) -
On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks
von: Stechly, Kaya, et al.
Veröffentlicht: (2024)