Towards Cost-Effective Reward Guided Text Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Rashid, Ahmad, Wu, Ruotian, Fan, Rongqi, Li, Hongliang, Kristiadi, Agustinus, Poupart, Pascal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Critical Look At Tokenwise Reward-Guided Text Generation
by: Rashid, Ahmad, et al.
Published: (2024)
by: Rashid, Ahmad, et al.
Published: (2024)
Uncertainty-Guided Likelihood Tree Search
by: Grosse, Julia, et al.
Published: (2024)
by: Grosse, Julia, et al.
Published: (2024)
Preventing Arbitrarily High Confidence on Far-Away Data in Point-Estimated Discriminative Neural Networks
by: Rashid, Ahmad, et al.
Published: (2023)
by: Rashid, Ahmad, et al.
Published: (2023)
Introduction to the Analysis of Probabilistic Decision-Making Algorithms
by: Kristiadi, Agustinus
Published: (2025)
by: Kristiadi, Agustinus
Published: (2025)
A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?
by: Kristiadi, Agustinus, et al.
Published: (2024)
by: Kristiadi, Agustinus, et al.
Published: (2024)
How Useful is Intermittent, Asynchronous Expert Feedback for Bayesian Optimization?
by: Kristiadi, Agustinus, et al.
Published: (2024)
by: Kristiadi, Agustinus, et al.
Published: (2024)
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
by: Cinquin, Tristan, et al.
Published: (2025)
by: Cinquin, Tristan, et al.
Published: (2025)
Time Is Effort: Estimating Human Post-Editing Time for Grammar Error Correction Tool Evaluation
by: Vadehra, Ankit, et al.
Published: (2025)
by: Vadehra, Ankit, et al.
Published: (2025)
On the Disconnect Between Theory and Practice of Neural Networks: Limits of the NTK Perspective
by: Wenger, Jonathan, et al.
Published: (2023)
by: Wenger, Jonathan, et al.
Published: (2023)
Reward-SQL: Boosting Text-to-SQL via Stepwise Reasoning and Process-Supervised Rewards
by: Zhang, Yuxin, et al.
Published: (2025)
by: Zhang, Yuxin, et al.
Published: (2025)
CERET: Cost-Effective Extrinsic Refinement for Text Generation
by: Cai, Jason, et al.
Published: (2024)
by: Cai, Jason, et al.
Published: (2024)
Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning
by: Chen, Yifang, et al.
Published: (2024)
by: Chen, Yifang, et al.
Published: (2024)
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization
by: Lu, Peng, et al.
Published: (2023)
by: Lu, Peng, et al.
Published: (2023)
Advancing Translation Preference Modeling with RLHF: A Step Towards Cost-Effective Solution
by: Xu, Nuo, et al.
Published: (2024)
by: Xu, Nuo, et al.
Published: (2024)
Balancing Cost and Effectiveness of Synthetic Data Generation Strategies for LLMs
by: Chan, Yung-Chieh, et al.
Published: (2024)
by: Chan, Yung-Chieh, et al.
Published: (2024)
Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms
by: Xiao, Zeguan, et al.
Published: (2025)
by: Xiao, Zeguan, et al.
Published: (2025)
FlashMD: long-stride, universal prediction of molecular dynamics
by: Bigi, Filippo, et al.
Published: (2025)
by: Bigi, Filippo, et al.
Published: (2025)
From Insight to Exploit: Leveraging LLM Collaboration for Adaptive Adversarial Text Generation
by: Sultana, Najrin, et al.
Published: (2025)
by: Sultana, Najrin, et al.
Published: (2025)
DreamReward: Text-to-3D Generation with Human Preference
by: Ye, Junliang, et al.
Published: (2024)
by: Ye, Junliang, et al.
Published: (2024)
Towards General-Purpose Text-Instruction-Guided Voice Conversion
by: Kuan, Chun-Yi, et al.
Published: (2023)
by: Kuan, Chun-Yi, et al.
Published: (2023)
ARGS: Alignment as Reward-Guided Search
by: Khanov, Maxim, et al.
Published: (2024)
by: Khanov, Maxim, et al.
Published: (2024)
ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
by: Zhang, Bonan, et al.
Published: (2025)
by: Zhang, Bonan, et al.
Published: (2025)
Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs
by: Liu, Hongliang, et al.
Published: (2026)
by: Liu, Hongliang, et al.
Published: (2026)
Low-Rank Filtering and Smoothing for Sequential Deep Learning
by: Sliwa, Joanna, et al.
Published: (2024)
by: Sliwa, Joanna, et al.
Published: (2024)
Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
by: Xie, Tianbao, et al.
Published: (2023)
by: Xie, Tianbao, et al.
Published: (2023)
Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework
by: Arias, Esteban Garces, et al.
Published: (2024)
by: Arias, Esteban Garces, et al.
Published: (2024)
Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts
by: Zhang, Di, et al.
Published: (2025)
by: Zhang, Di, et al.
Published: (2025)
GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation
by: Min, Do June, et al.
Published: (2024)
by: Min, Do June, et al.
Published: (2024)
Adaptive Contrastive Search: Uncertainty-Guided Decoding for Open-Ended Text Generation
by: Arias, Esteban Garces, et al.
Published: (2024)
by: Arias, Esteban Garces, et al.
Published: (2024)
On Designing Effective RL Reward at Training Time for LLM Reasoning
by: Gao, Jiaxuan, et al.
Published: (2024)
by: Gao, Jiaxuan, et al.
Published: (2024)
Mem-T: Densifying Rewards for Long-Horizon Memory Agents
by: Yue, Yanwei, et al.
Published: (2026)
by: Yue, Yanwei, et al.
Published: (2026)
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
Towards More Effective Table-to-Text Generation: Assessing In-Context Learning and Self-Evaluation with Open-Source Models
by: Iravani, Sahar, et al.
Published: (2024)
by: Iravani, Sahar, et al.
Published: (2024)
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
by: Tang, Xinyu, et al.
Published: (2025)
by: Tang, Xinyu, et al.
Published: (2025)
Review, Remask, Refine (R3): Process-Guided Block Diffusion for Text Generation
by: Mounier, Nikita, et al.
Published: (2025)
by: Mounier, Nikita, et al.
Published: (2025)
On-Policy RL with Optimal Reward Baseline
by: Hao, Yaru, et al.
Published: (2025)
by: Hao, Yaru, et al.
Published: (2025)
Fast-Slow Thinking RM: Efficient Integration of Scalar and Generative Reward Models
by: Wu, Jiayun, et al.
Published: (2026)
by: Wu, Jiayun, et al.
Published: (2026)
Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
by: Yang, Xinyi, et al.
Published: (2025)
by: Yang, Xinyi, et al.
Published: (2025)
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024)
by: Fisch, Adam, et al.
Published: (2024)
Similar Items
-
A Critical Look At Tokenwise Reward-Guided Text Generation
by: Rashid, Ahmad, et al.
Published: (2024) -
Uncertainty-Guided Likelihood Tree Search
by: Grosse, Julia, et al.
Published: (2024) -
Preventing Arbitrarily High Confidence on Far-Away Data in Point-Estimated Discriminative Neural Networks
by: Rashid, Ahmad, et al.
Published: (2023) -
Introduction to the Analysis of Probabilistic Decision-Making Algorithms
by: Kristiadi, Agustinus
Published: (2025) -
A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?
by: Kristiadi, Agustinus, et al.
Published: (2024)