LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Ang, Wang, Yifei, Yuan, Zhihang, Jegelka, Stefanie, Wang, Yisen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
von: Guo, Xiaojun, et al.
Veröffentlicht: (2025)
von: Guo, Xiaojun, et al.
Veröffentlicht: (2025)
How to Craft Backdoors with Unlabeled Data Alone?
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
When More is Less: Understanding Chain-of-Thought Length in LLMs
von: Wu, Yuyang, et al.
Veröffentlicht: (2025)
von: Wu, Yuyang, et al.
Veröffentlicht: (2025)
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
Understanding the Role of Equivariance in Self-supervised Learning
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
An Augmentation Overlap Theory of Contrastive Learning
von: Zhang, Qi, et al.
Veröffentlicht: (2025)
von: Zhang, Qi, et al.
Veröffentlicht: (2025)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
von: Wei, Zeming, et al.
Veröffentlicht: (2023)
von: Wei, Zeming, et al.
Veröffentlicht: (2023)
Dissecting the Failure of Invariant Learning on Graphs
von: Wang, Qixun, et al.
Veröffentlicht: (2024)
von: Wang, Qixun, et al.
Veröffentlicht: (2024)
Can In-context Learning Really Generalize to Out-of-distribution Tasks?
von: Wang, Qixun, et al.
Veröffentlicht: (2024)
von: Wang, Qixun, et al.
Veröffentlicht: (2024)
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
Do Generated Data Always Help Contrastive Learning?
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
A Canonicalization Perspective on Invariant and Equivariant Learning
von: Ma, George, et al.
Veröffentlicht: (2024)
von: Ma, George, et al.
Veröffentlicht: (2024)
Learning with Exact Invariances in Polynomial Time
von: Soleymani, Ashkan, et al.
Veröffentlicht: (2025)
von: Soleymani, Ashkan, et al.
Veröffentlicht: (2025)
Scaling Attention via Feature Sparsity
von: Xie, Yan, et al.
Veröffentlicht: (2026)
von: Xie, Yan, et al.
Veröffentlicht: (2026)
Bootstrapping Expectiles in Reinforcement Learning
von: Clavier, Pierre, et al.
Veröffentlicht: (2024)
von: Clavier, Pierre, et al.
Veröffentlicht: (2024)
Imitation Bootstrapped Reinforcement Learning
von: Hu, Hengyuan, et al.
Veröffentlicht: (2023)
von: Hu, Hengyuan, et al.
Veröffentlicht: (2023)
No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
von: Guo, Lixuan, et al.
Veröffentlicht: (2026)
von: Guo, Lixuan, et al.
Veröffentlicht: (2026)
Non-negative Contrastive Learning
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2025)
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2025)
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning
von: Zhong, Han, et al.
Veröffentlicht: (2025)
von: Zhong, Han, et al.
Veröffentlicht: (2025)
HRLAIF: Improvements in Helpfulness and Harmlessness in Open-domain Reinforcement Learning From AI Feedback
von: Li, Ang, et al.
Veröffentlicht: (2024)
von: Li, Ang, et al.
Veröffentlicht: (2024)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
von: Gatmiry, Khashayar, et al.
Veröffentlicht: (2024)
von: Gatmiry, Khashayar, et al.
Veröffentlicht: (2024)
Geometric Algorithms for Neural Combinatorial Optimization with Constraints
von: Karalias, Nikolaos, et al.
Veröffentlicht: (2025)
von: Karalias, Nikolaos, et al.
Veröffentlicht: (2025)
TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
Survey on Generalization Theory for Graph Neural Networks
von: Vasileiou, Antonis, et al.
Veröffentlicht: (2025)
von: Vasileiou, Antonis, et al.
Veröffentlicht: (2025)
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
von: Gupta, Sharut, et al.
Veröffentlicht: (2026)
von: Gupta, Sharut, et al.
Veröffentlicht: (2026)
VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
von: Wang, Yiting, et al.
Veröffentlicht: (2025)
von: Wang, Yiting, et al.
Veröffentlicht: (2025)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
A Theoretical Understanding of Self-Correction through In-context Alignment
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
RLAF: Reinforcement Learning from Automaton Feedback
von: Alinejad, Mahyar, et al.
Veröffentlicht: (2025)
von: Alinejad, Mahyar, et al.
Veröffentlicht: (2025)
Real-World Offline Reinforcement Learning from Vision Language Model Feedback
von: Venkataraman, Sreyas, et al.
Veröffentlicht: (2024)
von: Venkataraman, Sreyas, et al.
Veröffentlicht: (2024)
Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)
Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical Perspective
von: Zhang, Yi-Ge, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Ge, et al.
Veröffentlicht: (2025)
RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs
von: Wu, Jiaxing, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxing, et al.
Veröffentlicht: (2024)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
von: Wang, Xingyao, et al.
Veröffentlicht: (2023)
von: Wang, Xingyao, et al.
Veröffentlicht: (2023)
Enhancing LLMs for Physics Problem-Solving using Reinforcement Learning with Human-AI Feedback
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
Towards User-level Private Reinforcement Learning with Human Feedback
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
The Empirical Impact of Neural Parameter Symmetries, or Lack Thereof
von: Lim, Derek, et al.
Veröffentlicht: (2024)
von: Lim, Derek, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
von: Guo, Xiaojun, et al.
Veröffentlicht: (2025) -
How to Craft Backdoors with Unlabeled Data Alone?
von: Wang, Yifei, et al.
Veröffentlicht: (2024) -
When More is Less: Understanding Chain-of-Thought Length in LLMs
von: Wu, Yuyang, et al.
Veröffentlicht: (2025) -
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
von: Zhang, Qi, et al.
Veröffentlicht: (2024) -
Understanding the Role of Equivariance in Self-supervised Learning
von: Wang, Yifei, et al.
Veröffentlicht: (2024)