Credit Assignment with Resets in Language Model Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Samanta, Ankur, Magesh, Akshayaa, Jain, Ayush, Yu, Youliang, Jiang, Daniel, Asadi, Kavosh, Hassani, Kaveh, Sajda, Paul, Bhandari, Jalaj, Efroni, Yonathan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structure Enables Effective Self-Localization of Errors in LLMs
by: Samanta, Ankur, et al.
Published: (2026)
by: Samanta, Ankur, et al.
Published: (2026)
Self-Improvement of Language Models by Post-Training on Multi-Agent Debate
by: Samanta, Ankur, et al.
Published: (2025)
by: Samanta, Ankur, et al.
Published: (2025)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
by: Wu, Runzhe, et al.
Published: (2025)
by: Wu, Runzhe, et al.
Published: (2025)
Aligned Multi Objective Optimization
by: Efroni, Yonathan, et al.
Published: (2025)
by: Efroni, Yonathan, et al.
Published: (2025)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
by: Roth, Amit, et al.
Published: (2026)
by: Roth, Amit, et al.
Published: (2026)
Heterogeneous Multi-Player Multi-Armed Bandits Robust To Adversarial Attacks
by: Magesh, Akshayaa, et al.
Published: (2025)
by: Magesh, Akshayaa, et al.
Published: (2025)
Principled Detection of Hallucinations in Large Language Models via Multiple Testing
by: Li, Jiawei, et al.
Published: (2025)
by: Li, Jiawei, et al.
Published: (2025)
Gaze patterns predict preference and confidence in pairwise AI image evaluation
by: Papadopoulos, Nikolas, et al.
Published: (2026)
by: Papadopoulos, Nikolas, et al.
Published: (2026)
Gradient Free Deep Reinforcement Learning With TabPFN
by: Schiff, David, et al.
Published: (2025)
by: Schiff, David, et al.
Published: (2025)
Simple Optimizers for Convex Aligned Multi-Objective Optimization
by: Kretzu, Ben, et al.
Published: (2025)
by: Kretzu, Ben, et al.
Published: (2025)
The Bias of Harmful Label Associations in Vision-Language Models
by: Hazirbas, Caner, et al.
Published: (2024)
by: Hazirbas, Caner, et al.
Published: (2024)
Robust Multi-Hypothesis Testing with Moment Constrained Uncertainty Sets
by: Magesh, Akshayaa, et al.
Published: (2022)
by: Magesh, Akshayaa, et al.
Published: (2022)
Learning to Reason Efficiently with Discounted Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025)
by: Ayoub, Alex, et al.
Published: (2025)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
Adjoint sharding for very long context training of state space models
by: Xu, Xingzi, et al.
Published: (2025)
by: Xu, Xingzi, et al.
Published: (2025)
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
by: Pipano, Idan, et al.
Published: (2026)
by: Pipano, Idan, et al.
Published: (2026)
Exploiting Structure in Offline Multi-Agent RL: The Benefits of Low Interaction Rank
by: Zhan, Wenhao, et al.
Published: (2024)
by: Zhan, Wenhao, et al.
Published: (2024)
Pearl: A Production-ready Reinforcement Learning Agent
by: Zhu, Zheqing, et al.
Published: (2023)
by: Zhu, Zheqing, et al.
Published: (2023)
Aligning LLMs Toward Multi-Turn Conversational Outcomes Using Iterative PPO
by: Jiang, Daniel R., et al.
Published: (2025)
by: Jiang, Daniel R., et al.
Published: (2025)
From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models
by: Zhang, Chenchen
Published: (2026)
by: Zhang, Chenchen
Published: (2026)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
by: Parthasarathi, Prasanna, et al.
Published: (2025)
by: Parthasarathi, Prasanna, et al.
Published: (2025)
Learning the Target Network in Function Space
by: Asadi, Kavosh, et al.
Published: (2024)
by: Asadi, Kavosh, et al.
Published: (2024)
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Generalizing Multi-Step Inverse Models for Representation Learning to Finite-Memory POMDPs
by: Wu, Lili, et al.
Published: (2024)
by: Wu, Lili, et al.
Published: (2024)
Reducing Credit Assignment Variance via Counterfactual Reasoning Paths
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
by: Jiang, Xitai, et al.
Published: (2026)
by: Jiang, Xitai, et al.
Published: (2026)
Predicting Public Transportation Crowd Using Weather API and Social Media
by: Akshayaa Shree.M, Maheshwari.S
Published: (2026)
by: Akshayaa Shree.M, Maheshwari.S
Published: (2026)
The Generative Reasonable Person
by: Arbel, Yonathan A.
Published: (2025)
by: Arbel, Yonathan A.
Published: (2025)
Time After Time: Deep-Q Effect Estimation for Interventions on When and What to do
by: Wald, Yoav, et al.
Published: (2025)
by: Wald, Yoav, et al.
Published: (2025)
Ossarth : An Open-Source Customizable LLMOS
by: Magesh, Siddharth
Published: (2026)
by: Magesh, Siddharth
Published: (2026)
InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning
by: Yang, Matthew Y. R., et al.
Published: (2026)
by: Yang, Matthew Y. R., et al.
Published: (2026)
CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
by: Xie, Guofu, et al.
Published: (2025)
by: Xie, Guofu, et al.
Published: (2025)
Perception of an AI Teammate in an Embodied Control Task Affects Team Performance, Reflected in Human Teammates' Behaviors and Physiological Responses
by: Qin, Yinuo, et al.
Published: (2025)
by: Qin, Yinuo, et al.
Published: (2025)
Unleashing Dynamic Range and Resolution in Unlimited Sensing Framework via Novel Hardware
by: Zhu, Yuliang, et al.
Published: (2024)
by: Zhu, Yuliang, et al.
Published: (2024)
Unlimited Sampling of Multiband Signals: Single-Channel Acquisition and Recovery
by: Shtendel, Gal, et al.
Published: (2025)
by: Shtendel, Gal, et al.
Published: (2025)
1-Bit Unlimited Sampling Beyond Fourier Domain: Low-Resolution Sampling of Quantization Noise
by: Pavlicek, Vaclav, et al.
Published: (2025)
by: Pavlicek, Vaclav, et al.
Published: (2025)
USF Spectral Estimation: Prevalence of Gaussian Cramér-Rao Bounds Despite Modulo Folding
by: Guo, Ruiming, et al.
Published: (2025)
by: Guo, Ruiming, et al.
Published: (2025)
Sparse Sampling in Fractional Fourier Domain: Recovery Guarantees and Cramér-Rao Bounds
by: Pavlíček, Václav, et al.
Published: (2024)
by: Pavlíček, Václav, et al.
Published: (2024)
Blind Time-of-Flight Imaging: Sparse Deconvolution on the Continuum with Unknown Kernels
by: Guo, Ruiming, et al.
Published: (2024)
by: Guo, Ruiming, et al.
Published: (2024)
Unlocking Off-the-Grid Sparse Recovery with Unlimited Sensing: Simultaneous Super-Resolution in Time and Amplitude
by: Guo, Ruiming, et al.
Published: (2025)
by: Guo, Ruiming, et al.
Published: (2025)
Similar Items
-
Structure Enables Effective Self-Localization of Errors in LLMs
by: Samanta, Ankur, et al.
Published: (2026) -
Self-Improvement of Language Models by Post-Training on Multi-Agent Debate
by: Samanta, Ankur, et al.
Published: (2025) -
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
by: Wu, Runzhe, et al.
Published: (2025) -
Aligned Multi Objective Optimization
by: Efroni, Yonathan, et al.
Published: (2025) -
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
by: Roth, Amit, et al.
Published: (2026)