Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR
Fuente:
arXiv
Salvato in:
| Autori principali: | He, Yuhang, Wu, Haodong, Liu, Siyi, Ge, Hongyu, Zhou, Hange, Wu, Keyi, Zheng, Zhuo, Lin, Qihong, Zhong, Zixin, Zhang, Yongqi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hindsight Credit Assignment for Long-Horizon LLM Agents
di: Tan, Hui-Ze, et al.
Pubblicazione: (2026)
di: Tan, Hui-Ze, et al.
Pubblicazione: (2026)
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR
di: Kim, Soeun, et al.
Pubblicazione: (2026)
di: Kim, Soeun, et al.
Pubblicazione: (2026)
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
di: Lu, Jinda, et al.
Pubblicazione: (2026)
di: Lu, Jinda, et al.
Pubblicazione: (2026)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
di: Huang, Kexin, et al.
Pubblicazione: (2026)
di: Huang, Kexin, et al.
Pubblicazione: (2026)
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
di: Lu, Jinda, et al.
Pubblicazione: (2026)
di: Lu, Jinda, et al.
Pubblicazione: (2026)
Observational Constraints and Cosmological Dynamics of Interacting Fractional Holographic Dark Energy in Light of DESI DR2
di: Huang, Qihong, et al.
Pubblicazione: (2026)
di: Huang, Qihong, et al.
Pubblicazione: (2026)
StructSynth: Leveraging LLMs for Structure-Aware Tabular Data Synthesis in Low-Data Regimes
di: Liu, Siyi, et al.
Pubblicazione: (2025)
di: Liu, Siyi, et al.
Pubblicazione: (2025)
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR
di: Mou, Chaoli, et al.
Pubblicazione: (2026)
di: Mou, Chaoli, et al.
Pubblicazione: (2026)
From Instance Selection to Fixed-Pool Data Recipe Search for Supervised Fine-Tuning
di: Wu, Haodong, et al.
Pubblicazione: (2026)
di: Wu, Haodong, et al.
Pubblicazione: (2026)
Can Interlinked Credit and Insurance Contracts Boost Farmers' Adoption of Improved Rice Varieties? Field Experimental Evidence from China
di: Haixia Wu, et al.
Pubblicazione: (2025)
di: Haixia Wu, et al.
Pubblicazione: (2025)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2026)
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2026)
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
di: Yao, Jiashu, et al.
Pubblicazione: (2026)
di: Yao, Jiashu, et al.
Pubblicazione: (2026)
Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR
di: Min, Zijun, et al.
Pubblicazione: (2026)
di: Min, Zijun, et al.
Pubblicazione: (2026)
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
di: Wang, Jiakang, et al.
Pubblicazione: (2025)
di: Wang, Jiakang, et al.
Pubblicazione: (2025)
Mixture of Length and Pruning Experts for Knowledge Graphs Reasoning
di: Du, Enjun, et al.
Pubblicazione: (2025)
di: Du, Enjun, et al.
Pubblicazione: (2025)
GraphOracle: Efficient Fully-Inductive Knowledge Graph Reasoning via Relation-Dependency Graphs
di: Du, Enjun, et al.
Pubblicazione: (2025)
di: Du, Enjun, et al.
Pubblicazione: (2025)
Credit Where Credit is due: Recommendations for Inclusive and Ethical Authorship Practices in Ecology
di: Kevin C. Elliott, et al.
Pubblicazione: (2026)
di: Kevin C. Elliott, et al.
Pubblicazione: (2026)
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
di: Wang, Tao, et al.
Pubblicazione: (2026)
di: Wang, Tao, et al.
Pubblicazione: (2026)
Hindsight on High John
di: Moses, Richard
Pubblicazione: (1972)
di: Moses, Richard
Pubblicazione: (1972)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
di: Wu, Junkang, et al.
Pubblicazione: (2025)
di: Wu, Junkang, et al.
Pubblicazione: (2025)
Exploring Multi-Temperature Strategies for Token- and Rollout-Level Control in RLVR
di: Zhuang, Haomin, et al.
Pubblicazione: (2025)
di: Zhuang, Haomin, et al.
Pubblicazione: (2025)
Generalization of RLVR Using Causal Reasoning as a Testbed
di: Lu, Brian, et al.
Pubblicazione: (2025)
di: Lu, Brian, et al.
Pubblicazione: (2025)
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
di: Zhang, Kaiyi, et al.
Pubblicazione: (2026)
di: Zhang, Kaiyi, et al.
Pubblicazione: (2026)
Credit Where Credit Is Due: Considering Ethics, "Ethos", and Process in Library Instruction on Attribution
di: Harris, Benjamin R.
Pubblicazione: (2005)
di: Harris, Benjamin R.
Pubblicazione: (2005)
How Far Can Unsupervised RLVR Scale LLM Training?
di: He, Bingxiang, et al.
Pubblicazione: (2026)
di: He, Bingxiang, et al.
Pubblicazione: (2026)
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination
di: Zheng, Jiasheng, et al.
Pubblicazione: (2026)
di: Zheng, Jiasheng, et al.
Pubblicazione: (2026)
Mobile Cold Energy Storage: Coupling Food Distribution and Energy Systems
di: Lao, Hange, et al.
Pubblicazione: (2026)
di: Lao, Hange, et al.
Pubblicazione: (2026)
Self-Distilled RLVR
di: Yang, Chenxu, et al.
Pubblicazione: (2026)
di: Yang, Chenxu, et al.
Pubblicazione: (2026)
Tokens All the Way Down: A Money View of Decentralized Finance
di: Wu, Wenbin
Pubblicazione: (2026)
di: Wu, Wenbin
Pubblicazione: (2026)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
di: Wu, Yuning, et al.
Pubblicazione: (2026)
di: Wu, Yuning, et al.
Pubblicazione: (2026)
Generalized Capacity Planning for the Hospital-Residents Problem
di: Balasundaram, Haricharan, et al.
Pubblicazione: (2025)
di: Balasundaram, Haricharan, et al.
Pubblicazione: (2025)
From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering
di: Zhao, Yu, et al.
Pubblicazione: (2025)
di: Zhao, Yu, et al.
Pubblicazione: (2025)
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
di: Mitsuhashi, Ryo, et al.
Pubblicazione: (2026)
di: Mitsuhashi, Ryo, et al.
Pubblicazione: (2026)
RLVR-World: Training World Models with Reinforcement Learning
di: Wu, Jialong, et al.
Pubblicazione: (2025)
di: Wu, Jialong, et al.
Pubblicazione: (2025)
White Dwarfs with Infrared Excess from LAMOST Data Release 11
di: Wang, Keyi, et al.
Pubblicazione: (2026)
di: Wang, Keyi, et al.
Pubblicazione: (2026)
Leave or Stay? Community Environment Change and Resident Mobility in Tourism Gentrification
di: Zixin Feng, et al.
Pubblicazione: (2025)
di: Zixin Feng, et al.
Pubblicazione: (2025)
Translating Flow to Policy via Hindsight Online Imitation
di: Zheng, Yitian, et al.
Pubblicazione: (2025)
di: Zheng, Yitian, et al.
Pubblicazione: (2025)
Evaluating Parameter Efficient Methods for RLVR
di: Yin, Qingyu, et al.
Pubblicazione: (2025)
di: Yin, Qingyu, et al.
Pubblicazione: (2025)
Foreign Residency Rights and Corporate Bond Yield Spreads
di: Zhong‐qin Su, et al.
Pubblicazione: (2025)
di: Zhong‐qin Su, et al.
Pubblicazione: (2025)
Can Central Bank Digital Currencies Promote the Internationalization of Currencies?
di: Haodong Gu
Pubblicazione: (2025)
di: Haodong Gu
Pubblicazione: (2025)
Documenti analoghi
-
Hindsight Credit Assignment for Long-Horizon LLM Agents
di: Tan, Hui-Ze, et al.
Pubblicazione: (2026) -
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR
di: Kim, Soeun, et al.
Pubblicazione: (2026) -
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
di: Lu, Jinda, et al.
Pubblicazione: (2026) -
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
di: Huang, Kexin, et al.
Pubblicazione: (2026) -
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
di: Lu, Jinda, et al.
Pubblicazione: (2026)