RL for Reasoning by Adaptively Revealing Rationales
Fuente:
arXiv
Guardado en:
| Autores principales: | Amani, Mohammad Hossein, Lotfi, Aryo, Baldwin, Nicolas Mario, Bengio, Samy, Farajtabar, Mehrdad, Abbe, Emmanuel, West, Robert |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How Far Can Transformers Reason? The Globality Barrier and Inductive Scratchpad
por: Abbe, Emmanuel, et al.
Publicado: (2024)
por: Abbe, Emmanuel, et al.
Publicado: (2024)
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
por: Mahrooghi, Ilia, et al.
Publicado: (2026)
por: Mahrooghi, Ilia, et al.
Publicado: (2026)
Generalization on the Unseen, Logic Reasoning and Degree Curriculum
por: Abbe, Emmanuel, et al.
Publicado: (2023)
por: Abbe, Emmanuel, et al.
Publicado: (2023)
Chain-of-Sketch: Enabling Global Visual Reasoning
por: Lotfi, Aryo, et al.
Publicado: (2024)
por: Lotfi, Aryo, et al.
Publicado: (2024)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
por: Mirzadeh, Iman, et al.
Publicado: (2024)
por: Mirzadeh, Iman, et al.
Publicado: (2024)
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
por: Shojaee, Parshin, et al.
Publicado: (2025)
por: Shojaee, Parshin, et al.
Publicado: (2025)
Symbolic Autoencoding for Self-Supervised Sequence Learning
por: Amani, Mohammad Hossein, et al.
Publicado: (2024)
por: Amani, Mohammad Hossein, et al.
Publicado: (2024)
Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection
por: Chegini, Atoosa, et al.
Publicado: (2025)
por: Chegini, Atoosa, et al.
Publicado: (2025)
When can transformers reason with abstract symbols?
por: Boix-Adsera, Enric, et al.
Publicado: (2023)
por: Boix-Adsera, Enric, et al.
Publicado: (2023)
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
por: Gao, Silin, et al.
Publicado: (2025)
por: Gao, Silin, et al.
Publicado: (2025)
GRAD: Generative Retrieval-Aligned Demonstration Sampler for Efficient Few-Shot Reasoning
por: Gabouj, Oussama, et al.
Publicado: (2025)
por: Gabouj, Oussama, et al.
Publicado: (2025)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
por: Alizadeh, Keivan, et al.
Publicado: (2026)
por: Alizadeh, Keivan, et al.
Publicado: (2026)
Conformal Thinking: Risk Control for Reasoning on a Compute Budget
por: Wang, Xi, et al.
Publicado: (2026)
por: Wang, Xi, et al.
Publicado: (2026)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
por: Hannah, Lauren. A, et al.
Publicado: (2025)
por: Hannah, Lauren. A, et al.
Publicado: (2025)
$k$-server-bench: Automating Potential Discovery for the $k$-Server Conjecture
por: Brilliantov, Kirill, et al.
Publicado: (2026)
por: Brilliantov, Kirill, et al.
Publicado: (2026)
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
por: Joudaki, Amir, et al.
Publicado: (2025)
por: Joudaki, Amir, et al.
Publicado: (2025)
You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
por: Roy, Shuvendu, et al.
Publicado: (2025)
por: Roy, Shuvendu, et al.
Publicado: (2025)
Task Specific Sharpness Aware O-RAN Resource Management using Multi Agent Reinforcement Learning
por: Lotfi, Fatemeh, et al.
Publicado: (2025)
por: Lotfi, Fatemeh, et al.
Publicado: (2025)
ORAN-GUIDE: RAG-Driven Prompt Learning for LLM-Augmented Reinforcement Learning in O-RAN Network Slicing
por: Lotfi, Fatemeh, et al.
Publicado: (2025)
por: Lotfi, Fatemeh, et al.
Publicado: (2025)
Prompt-Tuned LLM-Augmented DRL for Dynamic O-RAN Network Slicing
por: Lotfi, Fatemeh, et al.
Publicado: (2025)
por: Lotfi, Fatemeh, et al.
Publicado: (2025)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
por: Mehta, Sachin, et al.
Publicado: (2024)
por: Mehta, Sachin, et al.
Publicado: (2024)
Solving Bayesian inverse problems with diffusion priors and off-policy RL
por: Scimeca, Luca, et al.
Publicado: (2025)
por: Scimeca, Luca, et al.
Publicado: (2025)
GFlowNet Foundations
por: Bengio, Yoshua, et al.
Publicado: (2021)
por: Bengio, Yoshua, et al.
Publicado: (2021)
GFlowNet Pretraining with Inexpensive Rewards
por: Pandey, Mohit, et al.
Publicado: (2024)
por: Pandey, Mohit, et al.
Publicado: (2024)
Bypassing the Rationale: Causal Auditing of Implicit Reasoning in Language Models
por: Sathyanarayanan, Anish, et al.
Publicado: (2026)
por: Sathyanarayanan, Anish, et al.
Publicado: (2026)
TIDE: Every Layer Knows the Token Beneath the Context
por: Jaiswal, Ajay, et al.
Publicado: (2026)
por: Jaiswal, Ajay, et al.
Publicado: (2026)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
por: Samragh, Mohammad, et al.
Publicado: (2024)
por: Samragh, Mohammad, et al.
Publicado: (2024)
SAFE-RL: Saliency-Aware Counterfactual Explainer for Deep Reinforcement Learning Policies
por: Samadi, Amir, et al.
Publicado: (2024)
por: Samadi, Amir, et al.
Publicado: (2024)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
por: Chen, Zihan, et al.
Publicado: (2025)
por: Chen, Zihan, et al.
Publicado: (2025)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
por: Maheswaran, Monishwaran, et al.
Publicado: (2025)
por: Maheswaran, Monishwaran, et al.
Publicado: (2025)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
por: Alizadeh, Keivan, et al.
Publicado: (2023)
por: Alizadeh, Keivan, et al.
Publicado: (2023)
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
por: Salimi, Moein, et al.
Publicado: (2026)
por: Salimi, Moein, et al.
Publicado: (2026)
Self-Evolving Curriculum for LLM Reasoning
por: Chen, Xiaoyin, et al.
Publicado: (2025)
por: Chen, Xiaoyin, et al.
Publicado: (2025)
Were RNNs All We Needed?
por: Feng, Leo, et al.
Publicado: (2024)
por: Feng, Leo, et al.
Publicado: (2024)
Structural Rationale Distillation via Reasoning Space Compression
por: Yang, Jialin, et al.
Publicado: (2026)
por: Yang, Jialin, et al.
Publicado: (2026)
ULLER: A Unified Language for Learning and Reasoning
por: van Krieken, Emile, et al.
Publicado: (2024)
por: van Krieken, Emile, et al.
Publicado: (2024)
FuzzingRL: Reinforcement Fuzz-Testing for Revealing VLM Failures
por: Xu, Jiajun, et al.
Publicado: (2026)
por: Xu, Jiajun, et al.
Publicado: (2026)
Token-Efficient RL for LLM Reasoning
por: Lee, Alan, et al.
Publicado: (2025)
por: Lee, Alan, et al.
Publicado: (2025)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
por: Armandpour, Mohammadreza, et al.
Publicado: (2026)
por: Armandpour, Mohammadreza, et al.
Publicado: (2026)
A Comprehensive Study of Supervised Machine Learning Models for Zero-Day Attack Detection: Analyzing Performance on Imbalanced Data
por: Lotfi, Zahra, et al.
Publicado: (2025)
por: Lotfi, Zahra, et al.
Publicado: (2025)
Ejemplares similares
-
How Far Can Transformers Reason? The Globality Barrier and Inductive Scratchpad
por: Abbe, Emmanuel, et al.
Publicado: (2024) -
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
por: Mahrooghi, Ilia, et al.
Publicado: (2026) -
Generalization on the Unseen, Logic Reasoning and Degree Curriculum
por: Abbe, Emmanuel, et al.
Publicado: (2023) -
Chain-of-Sketch: Enabling Global Visual Reasoning
por: Lotfi, Aryo, et al.
Publicado: (2024) -
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
por: Mirzadeh, Iman, et al.
Publicado: (2024)