Reinforcement Learning for LLM Post-Training: A Survey
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Zhichao, Ramnath, Kiran, Bi, Bin, Pentyala, Shiva Kumar, Chaudhuri, Sougata, Mehrotra, Shubham, Zixu, Zhu, Mao, Xiang-Bo, Asur, Sitaram, Na, Cheng |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PAFT: A Parallel Training Paradigm for Effective LLM Fine-Tuning
par: Pentyala, Shiva Kumar, et autres
Publié: (2024)
par: Pentyala, Shiva Kumar, et autres
Publié: (2024)
Rate, Explain and Cite (REC): Enhanced Explanation and Attribution in Automatic Evaluation by Large Language Models
par: Hsu, Aliyah R., et autres
Publié: (2024)
par: Hsu, Aliyah R., et autres
Publié: (2024)
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types
par: Wang, Zhichao, et autres
Publié: (2024)
par: Wang, Zhichao, et autres
Publié: (2024)
Diversity Enhances an LLM's Performance in RAG and Long-context Task
par: Wang, Zhichao, et autres
Publié: (2025)
par: Wang, Zhichao, et autres
Publié: (2025)
BayesFlow: A Probability Inference Framework for Meta-Agent Assisted Workflow Generation
par: Yuan, Bo, et autres
Publié: (2026)
par: Yuan, Bo, et autres
Publié: (2026)
Mutation invariants of cluster algebras of rank 2
par: Chen, Zhichao, et autres
Publié: (2024)
par: Chen, Zhichao, et autres
Publié: (2024)
A cluster theory approach from mutation invariants to Diophantine equations
par: Chen, Zhichao, et autres
Publié: (2025)
par: Chen, Zhichao, et autres
Publié: (2025)
UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function
par: Wang, Zhichao, et autres
Publié: (2024)
par: Wang, Zhichao, et autres
Publié: (2024)
A Quantitative Characterization of Forgetting in Post-Training
par: Balasubramanian, Krishnakumar, et autres
Publié: (2026)
par: Balasubramanian, Krishnakumar, et autres
Publié: (2026)
Reinforcement Mid-Training
par: Tian, Yijun, et autres
Publié: (2025)
par: Tian, Yijun, et autres
Publié: (2025)
Turning Conversations into Workflows: A Framework to Extract and Evaluate Dialog Workflows for Service AI Agents
par: Choubey, Prafulla Kumar, et autres
Publié: (2025)
par: Choubey, Prafulla Kumar, et autres
Publié: (2025)
The Role of Generator Access in Autoregressive Post-Training
par: Rege, Amit Kiran
Publié: (2026)
par: Rege, Amit Kiran
Publié: (2026)
Leveraging Data Symmetries to Select an Optimal Subset of Training Data under Label Noise
par: Shubham, Kumar, et autres
Publié: (2026)
par: Shubham, Kumar, et autres
Publié: (2026)
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
par: Hu, Pingbang, et autres
Publié: (2026)
par: Hu, Pingbang, et autres
Publié: (2026)
A Survey of Safe Reinforcement Learning and Constrained MDPs: A Technical Survey on Single-Agent and Multi-Agent Safety
par: Kushwaha, Ankita, et autres
Publié: (2025)
par: Kushwaha, Ankita, et autres
Publié: (2025)
Absence of a Putative Mannose-Specific Phosphotransferase System Enzyme IIAB Component in a Leucocin A-Resistant Strain ofListeria monocytogenes, as Shown by Two-Dimensional Sodium Dodecyl Sulfate-Polyacrylamide Gel Electrophoresis. / M. Ramnath
par: Ramnath, M
Publié: (2000)
par: Ramnath, M
Publié: (2000)
Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries
par: Ramnath, Sahana, et autres
Publié: (2025)
par: Ramnath, Sahana, et autres
Publié: (2025)
A multi-phase-field model for fiber-reinforced composite laminates based on puck failure theory
par: Kumar, Pavan Kumar Asur Vijaya, et autres
Publié: (2026)
par: Kumar, Pavan Kumar Asur Vijaya, et autres
Publié: (2026)
Alignment and Safety of Diffusion Models via Reinforcement Learning and Reward Modeling: A Survey
par: Lamba, Preeti, et autres
Publié: (2025)
par: Lamba, Preeti, et autres
Publié: (2025)
Learning to Ideate for Machine Learning Engineering Agents
par: Zhang, Yunxiang, et autres
Publié: (2026)
par: Zhang, Yunxiang, et autres
Publié: (2026)
Role-Based Fault Tolerance System for LLM RL Post-Training
par: Chen, Zhenqian, et autres
Publié: (2025)
par: Chen, Zhenqian, et autres
Publié: (2025)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
par: Rank, Ben, et autres
Publié: (2026)
par: Rank, Ben, et autres
Publié: (2026)
Two-strand ladder network variants: localization, multifractality, and quantum dynamics under an Aubry-André-Harper kind of quasiperiodicity
par: Biswas, Sougata
Publié: (2024)
par: Biswas, Sougata
Publié: (2024)
Topological phase transition and its stability against an applied magnetic field in a class of low dimensional decorated lattices
par: Biswas, Sougata
Publié: (2024)
par: Biswas, Sougata
Publié: (2024)
Portfolio management with the help of AI: What drives retail Indian investors to robo‐advisors?
par: Sougata Banerjee
Publié: (2024)
par: Sougata Banerjee
Publié: (2024)
Bayesimax Theory: Selecting Priors by Minimizing Total Information
par: Vangala, Sitaram
Publié: (2025)
par: Vangala, Sitaram
Publié: (2025)
CultureLLM: Incorporating Cultural Differences into Large Language Models
par: Li, Cheng, et autres
Publié: (2024)
par: Li, Cheng, et autres
Publié: (2024)
Brian Moore: An Ambassador of Feminism
par: Ramnath Singh Rathore
Publié: (2019)
par: Ramnath Singh Rathore
Publié: (2019)
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
par: Kumarappan, Adarsh, et autres
Publié: (2025)
par: Kumarappan, Adarsh, et autres
Publié: (2025)
Pivoting Retail Supply Chain with Deep Generative Techniques: Taxonomy, Survey and Insights
par: Wang, Yuan, et autres
Publié: (2024)
par: Wang, Yuan, et autres
Publié: (2024)
Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models
par: Singh, Mukul, et autres
Publié: (2025)
par: Singh, Mukul, et autres
Publié: (2025)
FeedbackLLM: Metadata driven Multi-Agentic Language Agnostic Test Case Generator with Evolving prompt and Coverage Feedback
par: Jasti, Kushal, et autres
Publié: (2026)
par: Jasti, Kushal, et autres
Publié: (2026)
InsightBuild: LLM-Powered Causal Reasoning in Smart Building Systems
par: Neogi, Pinaki Prasad Guha, et autres
Publié: (2025)
par: Neogi, Pinaki Prasad Guha, et autres
Publié: (2025)
Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning
par: Zhang, Lingzhe, et autres
Publié: (2026)
par: Zhang, Lingzhe, et autres
Publié: (2026)
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
par: Tan, Zelin, et autres
Publié: (2025)
par: Tan, Zelin, et autres
Publié: (2025)
LATENT: LLM-Augmented Trojan Insertion and Evaluation Framework for Analog Netlist Topologies
par: Chaudhuri, Jayeeta, et autres
Publié: (2025)
par: Chaudhuri, Jayeeta, et autres
Publié: (2025)
FastLane: Efficient Routed Systems for Late-Interaction Retrieval
par: Kumar, Ramnath, et autres
Publié: (2026)
par: Kumar, Ramnath, et autres
Publié: (2026)
Secure Cross-Silo Synthetic Genomic Data Generation
par: Filienko, Daniil, et autres
Publié: (2026)
par: Filienko, Daniil, et autres
Publié: (2026)
Numerical study of unsteady bioconvective transport of oxytactic microorganisms over a stretching cone
par: Pentyala Srinivasa Rao, et autres
Publié: (2025)
par: Pentyala Srinivasa Rao, et autres
Publié: (2025)
CaPS: Collaborative and Private Synthetic Data Generation from Distributed Sources
par: Pentyala, Sikha, et autres
Publié: (2024)
par: Pentyala, Sikha, et autres
Publié: (2024)
Documents similaires
-
PAFT: A Parallel Training Paradigm for Effective LLM Fine-Tuning
par: Pentyala, Shiva Kumar, et autres
Publié: (2024) -
Rate, Explain and Cite (REC): Enhanced Explanation and Attribution in Automatic Evaluation by Large Language Models
par: Hsu, Aliyah R., et autres
Publié: (2024) -
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types
par: Wang, Zhichao, et autres
Publié: (2024) -
Diversity Enhances an LLM's Performance in RAG and Long-context Task
par: Wang, Zhichao, et autres
Publié: (2025) -
BayesFlow: A Probability Inference Framework for Meta-Agent Assisted Workflow Generation
par: Yuan, Bo, et autres
Publié: (2026)