Exploiting Tree Structure for Credit Assignment in RL Training of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Tran, Hieu, Yao, Zonghai, Yu, Hong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VinePPO: Refining Credit Assignment in RL Training of LLMs
by: Kazemnejad, Amirhossein, et al.
Published: (2024)
by: Kazemnejad, Amirhossein, et al.
Published: (2024)
BioInstruct: Instruction Tuning of Large Language Models for Biomedical Natural Language Processing
by: Tran, Hieu, et al.
Published: (2023)
by: Tran, Hieu, et al.
Published: (2023)
ReadCtrl: Personalizing text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2024)
by: Tran, Hieu, et al.
Published: (2024)
Enhancing LLMs for Identifying and Prioritizing Important Medical Jargons from Electronic Health Record Notes Utilizing Data Augmentation
by: Jang, Won Seok, et al.
Published: (2025)
by: Jang, Won Seok, et al.
Published: (2025)
RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models
by: Tran, Hieu, et al.
Published: (2024)
by: Tran, Hieu, et al.
Published: (2024)
PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning
by: Tran, Hieu, et al.
Published: (2025)
by: Tran, Hieu, et al.
Published: (2025)
MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2025)
by: Tran, Hieu, et al.
Published: (2025)
JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability
by: Wang, Junda, et al.
Published: (2024)
by: Wang, Junda, et al.
Published: (2024)
EHR Interaction Between Patients and AI: NoteAid EHR Interaction
by: Zhang, Xiaocheng, et al.
Published: (2023)
by: Zhang, Xiaocheng, et al.
Published: (2023)
Large Language Models are In-context Teachers for Knowledge Reasoning
by: Zhao, Jiachen, et al.
Published: (2023)
by: Zhao, Jiachen, et al.
Published: (2023)
Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models
by: Guo, Yiran, et al.
Published: (2025)
by: Guo, Yiran, et al.
Published: (2025)
README: Bridging Medical Jargon and Lay Understanding for Patient Education through Data-Centric NLP
by: Yao, Zonghai, et al.
Published: (2023)
by: Yao, Zonghai, et al.
Published: (2023)
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
by: Machcha, Sravanthi, et al.
Published: (2026)
by: Machcha, Sravanthi, et al.
Published: (2026)
Efficient and Effective Internal Memory Retrieval for LLM-Based Healthcare Prediction
by: Li, Mingchen, et al.
Published: (2026)
by: Li, Mingchen, et al.
Published: (2026)
Beyond Uniform Credit: Causal Credit Assignment for Policy Optimization
by: Khandoga, Mykola, et al.
Published: (2026)
by: Khandoga, Mykola, et al.
Published: (2026)
CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
by: Xie, Guofu, et al.
Published: (2025)
by: Xie, Guofu, et al.
Published: (2025)
Do Physicians Know How to Prompt? The Need for Automatic Prompt Optimization Help in Clinical Note Generation
by: Yao, Zonghai, et al.
Published: (2023)
by: Yao, Zonghai, et al.
Published: (2023)
Large Language Model-based Role-Playing for Personalized Medical Jargon Extraction
by: Lim, Jung Hoon, et al.
Published: (2024)
by: Lim, Jung Hoon, et al.
Published: (2024)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
by: Phan, Phuc, et al.
Published: (2024)
by: Phan, Phuc, et al.
Published: (2024)
Blocks Architecture (BloArk): Efficient, Cost-Effective, and Incremental Dataset Architecture for Wikipedia Revision History
by: Li, Lingxi, et al.
Published: (2024)
by: Li, Lingxi, et al.
Published: (2024)
From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models
by: Zhang, Chenchen
Published: (2026)
by: Zhang, Chenchen
Published: (2026)
Advancing Language Multi-Agent Learning with Credit Re-Assignment for Interactive Environment Generalization
by: He, Zhitao, et al.
Published: (2025)
by: He, Zhitao, et al.
Published: (2025)
Every Step Counts: Step-Level Credit Assignment for Tool-Integrated Text-to-SQL
by: Dai, Yaxun, et al.
Published: (2026)
by: Dai, Yaxun, et al.
Published: (2026)
MedCOD: Enhancing English-to-Spanish Medical Translation of Large Language Models Using Enriched Chain-of-Dictionary Framework
by: Salim, Md Shahidul, et al.
Published: (2025)
by: Salim, Md Shahidul, et al.
Published: (2025)
Exploiting LLMs' Reasoning Capability to Infer Implicit Concepts in Legal Information Retrieval
by: Nguyen, Hai-Long, et al.
Published: (2024)
by: Nguyen, Hai-Long, et al.
Published: (2024)
Reducing Credit Assignment Variance via Counterfactual Reasoning Paths
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic
by: Zhang, Yaocheng, et al.
Published: (2025)
by: Zhang, Yaocheng, et al.
Published: (2025)
When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
by: Wang, Jiakang, et al.
Published: (2025)
by: Wang, Jiakang, et al.
Published: (2025)
Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use
by: Kumar, Abhijit, et al.
Published: (2026)
by: Kumar, Abhijit, et al.
Published: (2026)
APEX-Searcher: Refining Credit Assignment with Subgoaling for Agentic Retrieval-Augmented Generation
by: Chen, Kun, et al.
Published: (2026)
by: Chen, Kun, et al.
Published: (2026)
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
by: Li, Ziheng, et al.
Published: (2026)
by: Li, Ziheng, et al.
Published: (2026)
Exploration Hacking: Can LLMs Learn to Resist RL Training?
by: Jang, Eyon, et al.
Published: (2026)
by: Jang, Eyon, et al.
Published: (2026)
Rejection Improves Reliability: Training LLMs to Refuse Unknown Questions Using RL from Knowledge Feedback
by: Xu, Hongshen, et al.
Published: (2024)
by: Xu, Hongshen, et al.
Published: (2024)
DischargeSim: A Simulation Benchmark for Educational Doctor-Patient Communication at Discharge
by: Yao, Zonghai, et al.
Published: (2025)
by: Yao, Zonghai, et al.
Published: (2025)
InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning
by: Yang, Matthew Y. R., et al.
Published: (2026)
by: Yang, Matthew Y. R., et al.
Published: (2026)
Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance
by: Chen, Xinzhu, et al.
Published: (2026)
by: Chen, Xinzhu, et al.
Published: (2026)
Teaching LLMs at Charles University: Assignments and Activities
by: Helcl, Jindřich, et al.
Published: (2024)
by: Helcl, Jindřich, et al.
Published: (2024)
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
by: Zhang, Kaiyi, et al.
Published: (2026)
by: Zhang, Kaiyi, et al.
Published: (2026)
MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
by: Zhang, Naifan, et al.
Published: (2026)
by: Zhang, Naifan, et al.
Published: (2026)
Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers
by: Hu, Senkang, et al.
Published: (2026)
by: Hu, Senkang, et al.
Published: (2026)
Similar Items
-
VinePPO: Refining Credit Assignment in RL Training of LLMs
by: Kazemnejad, Amirhossein, et al.
Published: (2024) -
BioInstruct: Instruction Tuning of Large Language Models for Biomedical Natural Language Processing
by: Tran, Hieu, et al.
Published: (2023) -
ReadCtrl: Personalizing text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2024) -
Enhancing LLMs for Identifying and Prioritizing Important Medical Jargons from Electronic Health Record Notes Utilizing Data Augmentation
by: Jang, Won Seok, et al.
Published: (2025) -
RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models
by: Tran, Hieu, et al.
Published: (2024)