Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Haikang, Raffel, Colin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tool-Augmented Reward Modeling
von: Li, Lei, et al.
Veröffentlicht: (2023)
von: Li, Lei, et al.
Veröffentlicht: (2023)
Decoupling Task-Solving and Output Formatting in LLM Generation
von: Deng, Haikang, et al.
Veröffentlicht: (2025)
von: Deng, Haikang, et al.
Veröffentlicht: (2025)
RAGferee: Building Contextual Reward Models for Retrieval-Augmented Generation
von: Coman, Andrei C., et al.
Veröffentlicht: (2025)
von: Coman, Andrei C., et al.
Veröffentlicht: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
Cascade Reward Sampling for Efficient Decoding-Time Alignment
von: Li, Bolian, et al.
Veröffentlicht: (2024)
von: Li, Bolian, et al.
Veröffentlicht: (2024)
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
von: Jin, Zhuoran, et al.
Veröffentlicht: (2024)
von: Jin, Zhuoran, et al.
Veröffentlicht: (2024)
GRAM: A Generative Foundation Reward Model for Reward Generalization
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
On the Low-Rank Parametrization of Reward Models for Controlled Language Generation
von: Troshin, Sergey, et al.
Veröffentlicht: (2024)
von: Troshin, Sergey, et al.
Veröffentlicht: (2024)
Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
von: Hu, Yushi, et al.
Veröffentlicht: (2025)
von: Hu, Yushi, et al.
Veröffentlicht: (2025)
Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
von: Xu, Zhichao, et al.
Veröffentlicht: (2025)
von: Xu, Zhichao, et al.
Veröffentlicht: (2025)
Merging by Matching Models in Task Parameter Subspaces
von: Tam, Derek, et al.
Veröffentlicht: (2023)
von: Tam, Derek, et al.
Veröffentlicht: (2023)
Position: The Most Expensive Part of an LLM should be its Training Data
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2025)
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2025)
Long-form RewardBench: Evaluating Reward Models for Long-form Generation
von: Huang, Hui, et al.
Veröffentlicht: (2026)
von: Huang, Hui, et al.
Veröffentlicht: (2026)
Controlling Multimodal LLMs via Reward-guided Decoding
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
Training Language Models to Generate Text with Citations via Fine-grained Rewards
von: Huang, Chengyu, et al.
Veröffentlicht: (2024)
von: Huang, Chengyu, et al.
Veröffentlicht: (2024)
Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Models
von: Chen, Yihang, et al.
Veröffentlicht: (2025)
von: Chen, Yihang, et al.
Veröffentlicht: (2025)
Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
von: Xie, Tianbao, et al.
Veröffentlicht: (2023)
von: Xie, Tianbao, et al.
Veröffentlicht: (2023)
Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
von: Elle
Veröffentlicht: (2025)
von: Elle
Veröffentlicht: (2025)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024)
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024)
Towards Cost-Effective Reward Guided Text Generation
von: Rashid, Ahmad, et al.
Veröffentlicht: (2025)
von: Rashid, Ahmad, et al.
Veröffentlicht: (2025)
HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
RewardBench 2: Advancing Reward Model Evaluation
von: Malik, Saumya, et al.
Veröffentlicht: (2025)
von: Malik, Saumya, et al.
Veröffentlicht: (2025)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards
von: Pisano, Raffaele, et al.
Veröffentlicht: (2026)
von: Pisano, Raffaele, et al.
Veröffentlicht: (2026)
Reward Models Can Improve Themselves: Reward-Guided Adversarial Failure Mode Discovery for Robust Reward Modeling
von: Pathmanathan, Pankayaraj, et al.
Veröffentlicht: (2025)
von: Pathmanathan, Pankayaraj, et al.
Veröffentlicht: (2025)
Decoding-Time Debiasing via Process Reward Models: From Controlled Fill-in to Open-Ended Generation
von: Khan, Muneeb Ur Raheem
Veröffentlicht: (2026)
von: Khan, Muneeb Ur Raheem
Veröffentlicht: (2026)
RRM: Robust Reward Model Training Mitigates Reward Hacking
von: Liu, Tianqi, et al.
Veröffentlicht: (2024)
von: Liu, Tianqi, et al.
Veröffentlicht: (2024)
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
von: Liu, Runheng, et al.
Veröffentlicht: (2026)
von: Liu, Runheng, et al.
Veröffentlicht: (2026)
RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards
von: Li, Xinze, et al.
Veröffentlicht: (2024)
von: Li, Xinze, et al.
Veröffentlicht: (2024)
SeRTS: Self-Rewarding Tree Search for Biomedical Retrieval-Augmented Generation
von: Hu, Minda, et al.
Veröffentlicht: (2024)
von: Hu, Minda, et al.
Veröffentlicht: (2024)
A Critical Look At Tokenwise Reward-Guided Text Generation
von: Rashid, Ahmad, et al.
Veröffentlicht: (2024)
von: Rashid, Ahmad, et al.
Veröffentlicht: (2024)
Reward Reasoning Model
von: Guo, Jiaxin, et al.
Veröffentlicht: (2025)
von: Guo, Jiaxin, et al.
Veröffentlicht: (2025)
Dynamic Multi-Reward Weighting for Multi-Style Controllable Generation
von: de Langis, Karin, et al.
Veröffentlicht: (2024)
von: de Langis, Karin, et al.
Veröffentlicht: (2024)
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
Reward-SQL: Boosting Text-to-SQL via Stepwise Reasoning and Process-Supervised Rewards
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
Realistic Evaluation of Model Merging for Compositional Generalization
von: Tam, Derek, et al.
Veröffentlicht: (2024)
von: Tam, Derek, et al.
Veröffentlicht: (2024)
GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
Reason Only When Needed: Efficient Generative Reward Modeling via Model-Internal Uncertainty
von: Xue, Chao, et al.
Veröffentlicht: (2026)
von: Xue, Chao, et al.
Veröffentlicht: (2026)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Tool-Augmented Reward Modeling
von: Li, Lei, et al.
Veröffentlicht: (2023) -
Decoupling Task-Solving and Output Formatting in LLM Generation
von: Deng, Haikang, et al.
Veröffentlicht: (2025) -
RAGferee: Building Contextual Reward Models for Retrieval-Augmented Generation
von: Coman, Andrei C., et al.
Veröffentlicht: (2025) -
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025) -
Cascade Reward Sampling for Efficient Decoding-Time Alignment
von: Li, Bolian, et al.
Veröffentlicht: (2024)