RL in Name Only? Analyzing the Structural Assumptions in RL post-training for LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Samineni, Soumya Rani, Kalwar, Durgesh, Valmeekam, Karthik, Stechly, Kaya, Kambhampati, Subbarao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
por: Samineni, Soumya Rani, et al.
Publicado: (2025)
por: Samineni, Soumya Rani, et al.
Publicado: (2025)
Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
por: Kambhampati, Subbarao, et al.
Publicado: (2025)
por: Kambhampati, Subbarao, et al.
Publicado: (2025)
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
por: Valmeekam, Karthik, et al.
Publicado: (2024)
por: Valmeekam, Karthik, et al.
Publicado: (2024)
Chain of Thoughtlessness? An Analysis of CoT in Planning
por: Stechly, Kaya, et al.
Publicado: (2024)
por: Stechly, Kaya, et al.
Publicado: (2024)
On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks
por: Stechly, Kaya, et al.
Publicado: (2024)
por: Stechly, Kaya, et al.
Publicado: (2024)
Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
por: Valmeekam, Karthik, et al.
Publicado: (2025)
por: Valmeekam, Karthik, et al.
Publicado: (2025)
Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
por: Palod, Vardhan, et al.
Publicado: (2025)
por: Palod, Vardhan, et al.
Publicado: (2025)
Planning in Strawberry Fields: Evaluating and Improving the Planning and Scheduling Capabilities of LRM o1
por: Valmeekam, Karthik, et al.
Publicado: (2024)
por: Valmeekam, Karthik, et al.
Publicado: (2024)
LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
por: Kambhampati, Subbarao, et al.
Publicado: (2024)
por: Kambhampati, Subbarao, et al.
Publicado: (2024)
(How) Do reasoning models reason?
por: Subbarao Kambhampati, et al.
Publicado: (2025)
por: Subbarao Kambhampati, et al.
Publicado: (2025)
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning
por: Bhambri, Siddhant, et al.
Publicado: (2024)
por: Bhambri, Siddhant, et al.
Publicado: (2024)
Can Large Language Models Reason and Plan?
por: Kambhampati, Subbarao
Publicado: (2024)
por: Kambhampati, Subbarao
Publicado: (2024)
Mind The Gap: Quantifying Mechanistic Gaps in Algorithmic Reasoning via Neural Compilation
por: Saldyt, Lucas, et al.
Publicado: (2025)
por: Saldyt, Lucas, et al.
Publicado: (2025)
Efficiency Boost in Decentralized Optimization: Reimagining Neighborhood Aggregation with Minimal Overhead
por: Kalwar, Durgesh, et al.
Publicado: (2025)
por: Kalwar, Durgesh, et al.
Publicado: (2025)
Robust Planning with Compound LLM Architectures: An LLM-Modulo Approach
por: Gundawar, Atharva, et al.
Publicado: (2024)
por: Gundawar, Atharva, et al.
Publicado: (2024)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
por: Xiao, Wei, et al.
Publicado: (2025)
por: Xiao, Wei, et al.
Publicado: (2025)
Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning
por: Gundawar, Atharva, et al.
Publicado: (2024)
por: Gundawar, Atharva, et al.
Publicado: (2024)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
por: Zhang, Yiqi, et al.
Publicado: (2026)
por: Zhang, Yiqi, et al.
Publicado: (2026)
Partial Policy Gradients for RL in LLMs
por: Mathur, Puneet, et al.
Publicado: (2026)
por: Mathur, Puneet, et al.
Publicado: (2026)
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
por: Ran-Milo, Yuval, et al.
Publicado: (2026)
por: Ran-Milo, Yuval, et al.
Publicado: (2026)
A Systematic Investigation of The RL-Jailbreaker in LLMs
por: Mohammedalamen, Montaser, et al.
Publicado: (2026)
por: Mohammedalamen, Montaser, et al.
Publicado: (2026)
Prioritized Replay for RL Post-training
por: Fatemi, Mehdi
Publicado: (2026)
por: Fatemi, Mehdi
Publicado: (2026)
RL$^3$: Boosting Meta Reinforcement Learning via RL inside RL$^2$
por: Bhatia, Abhinav, et al.
Publicado: (2023)
por: Bhatia, Abhinav, et al.
Publicado: (2023)
RL-Struct: A Lightweight Reinforcement Learning Framework for Reliable Structured Output in LLMs
por: Hu, Ruike, et al.
Publicado: (2025)
por: Hu, Ruike, et al.
Publicado: (2025)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
por: Mark, Max Sobol, et al.
Publicado: (2024)
por: Mark, Max Sobol, et al.
Publicado: (2024)
Toward Explainable Offline RL: Analyzing Representations in Intrinsically Motivated Decision Transformers
por: Guiducci, Leonardo, et al.
Publicado: (2025)
por: Guiducci, Leonardo, et al.
Publicado: (2025)
Incorporating Human Flexibility through Reward Preferences in Human-AI Teaming
por: Bhambri, Siddhant, et al.
Publicado: (2023)
por: Bhambri, Siddhant, et al.
Publicado: (2023)
A Comedy of Estimators: On KL Regularization in RL Training of LLMs
por: Shah, Vedant, et al.
Publicado: (2025)
por: Shah, Vedant, et al.
Publicado: (2025)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
por: Wu, Runzhe, et al.
Publicado: (2025)
por: Wu, Runzhe, et al.
Publicado: (2025)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
por: Pignatelli, Eduardo, et al.
Publicado: (2024)
por: Pignatelli, Eduardo, et al.
Publicado: (2024)
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2026)
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2026)
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
por: Huang, Luke J., et al.
Publicado: (2026)
por: Huang, Luke J., et al.
Publicado: (2026)
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
por: Su, Jianhai, et al.
Publicado: (2025)
por: Su, Jianhai, et al.
Publicado: (2025)
Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward
por: Wen, Xuexiang, et al.
Publicado: (2026)
por: Wen, Xuexiang, et al.
Publicado: (2026)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
por: Chen, Zihan, et al.
Publicado: (2025)
por: Chen, Zihan, et al.
Publicado: (2025)
Combining LLM decision and RL action selection to improve RL policy for adaptive interventions
por: Karine, Karine, et al.
Publicado: (2025)
por: Karine, Karine, et al.
Publicado: (2025)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
por: Muslimani, Calarina, et al.
Publicado: (2025)
por: Muslimani, Calarina, et al.
Publicado: (2025)
RL-Finetuned LLMs for Privacy-Preserving Synthetic Rewriting
por: Shi, Zhan, et al.
Publicado: (2025)
por: Shi, Zhan, et al.
Publicado: (2025)
EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation
por: Shi, Jiahe, et al.
Publicado: (2025)
por: Shi, Jiahe, et al.
Publicado: (2025)
Position Paper: Rethinking Privacy in RL for Sequential Decision-making in the Age of LLMs
por: Fan, Flint Xiaofeng, et al.
Publicado: (2025)
por: Fan, Flint Xiaofeng, et al.
Publicado: (2025)
Ejemplares similares
-
Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
por: Samineni, Soumya Rani, et al.
Publicado: (2025) -
Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
por: Kambhampati, Subbarao, et al.
Publicado: (2025) -
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
por: Valmeekam, Karthik, et al.
Publicado: (2024) -
Chain of Thoughtlessness? An Analysis of CoT in Planning
por: Stechly, Kaya, et al.
Publicado: (2024) -
On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks
por: Stechly, Kaya, et al.
Publicado: (2024)