Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yifang, Wang, Shuohang, Yang, Ziyi, Sharma, Hiteshi, Karampatziakis, Nikos, Yu, Donghan, Jamieson, Kevin, Du, Simon Shaolei, Shen, Yelong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Temperature-Centric Investigation of Speculative Decoding with Knowledge Distillation
by: Ouyang, Siru, et al.
Published: (2024)
by: Ouyang, Siru, et al.
Published: (2024)
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
by: Wang, Yiping, et al.
Published: (2024)
by: Wang, Yiping, et al.
Published: (2024)
Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Rethinking Data Synthesis: A Teacher Model Training Recipe with Interpretation
by: Chen, Yifang, et al.
Published: (2024)
by: Chen, Yifang, et al.
Published: (2024)
Multi-LoRA Composition for Image Generation
by: Zhong, Ming, et al.
Published: (2024)
by: Zhong, Ming, et al.
Published: (2024)
Variance Alignment Score: A Simple But Tough-to-Beat Data Selection Method for Multimodal Contrastive Learning
by: Wang, Yiping, et al.
Published: (2024)
by: Wang, Yiping, et al.
Published: (2024)
Active, anytime-valid risk controlling prediction sets
by: Xu, Ziyu, et al.
Published: (2024)
by: Xu, Ziyu, et al.
Published: (2024)
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
by: Wang, Yiping, et al.
Published: (2025)
by: Wang, Yiping, et al.
Published: (2025)
ThetaEvolve: Test-time Learning on Open Problems
by: Wang, Yiping, et al.
Published: (2025)
by: Wang, Yiping, et al.
Published: (2025)
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
by: Zhang, Zhen, et al.
Published: (2025)
by: Zhang, Zhen, et al.
Published: (2025)
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
by: Zhang, Rongzhi, et al.
Published: (2024)
by: Zhang, Rongzhi, et al.
Published: (2024)
CLIPLoss and Norm-Based Data Selection Methods for Multimodal Contrastive Learning
by: Wang, Yiping, et al.
Published: (2024)
by: Wang, Yiping, et al.
Published: (2024)
LoRe: Personalizing LLMs via Low-Rank Reward Modeling
by: Bose, Avinandan, et al.
Published: (2025)
by: Bose, Avinandan, et al.
Published: (2025)
Towards Cost-Effective Reward Guided Text Generation
by: Rashid, Ahmad, et al.
Published: (2025)
by: Rashid, Ahmad, et al.
Published: (2025)
Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval
by: Li, Siting, et al.
Published: (2025)
by: Li, Siting, et al.
Published: (2025)
qa-FLoRA: Data-free query-adaptive Fusion of LoRAs for LLMs
by: Shukla, Shreya, et al.
Published: (2025)
by: Shukla, Shreya, et al.
Published: (2025)
Mojito: Motion Trajectory and Intensity Control for Video Generation
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively
by: Gu, Jiawei, et al.
Published: (2025)
by: Gu, Jiawei, et al.
Published: (2025)
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
by: Li, Siting, et al.
Published: (2024)
by: Li, Siting, et al.
Published: (2024)
Rethinking the Role of Proxy Rewards in Language Model Alignment
by: Kim, Sungdong, et al.
Published: (2024)
by: Kim, Sungdong, et al.
Published: (2024)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
by: Huang, Zeyi, et al.
Published: (2026)
by: Huang, Zeyi, et al.
Published: (2026)
Large Language Models Are Read/Write Policy-Makers for Simultaneous Generation
by: Guo, Shoutao, et al.
Published: (2025)
by: Guo, Shoutao, et al.
Published: (2025)
Transformers are Efficient Compilers, Provably
by: Zhai, Xiyu, et al.
Published: (2024)
by: Zhai, Xiyu, et al.
Published: (2024)
Policy-Based Trajectory Clustering in Offline Reinforcement Learning
by: Hu, Hao, et al.
Published: (2025)
by: Hu, Hao, et al.
Published: (2025)
Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards
by: Ren, Mengjie, et al.
Published: (2026)
by: Ren, Mengjie, et al.
Published: (2026)
Reinforcement World Model Learning for LLM-based Agents
by: Yu, Xiao, et al.
Published: (2026)
by: Yu, Xiao, et al.
Published: (2026)
Auto-RAG: Autonomous Retrieval-Augmented Generation for Large Language Models
by: Yu, Tian, et al.
Published: (2024)
by: Yu, Tian, et al.
Published: (2024)
Truth-Aware Context Selection: Mitigating Hallucinations of Large Language Models Being Misled by Untruthful Contexts
by: Yu, Tian, et al.
Published: (2024)
by: Yu, Tian, et al.
Published: (2024)
Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math
by: Xu, Haoran, et al.
Published: (2025)
by: Xu, Haoran, et al.
Published: (2025)
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
by: Bai, Yang, et al.
Published: (2026)
by: Bai, Yang, et al.
Published: (2026)
Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models
by: Hong, Haitao, et al.
Published: (2025)
by: Hong, Haitao, et al.
Published: (2025)
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
by: Liu, Mickel, et al.
Published: (2025)
by: Liu, Mickel, et al.
Published: (2025)
Reviving The Classics: Active Reward Modeling in Large Language Model Alignment
by: Shen, Yunyi, et al.
Published: (2025)
by: Shen, Yunyi, et al.
Published: (2025)
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
by: Ye, Ziyi, et al.
Published: (2024)
by: Ye, Ziyi, et al.
Published: (2024)
Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
by: Ren, Liliang, et al.
Published: (2025)
by: Ren, Liliang, et al.
Published: (2025)
An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models
by: Bhatt, Gantavya, et al.
Published: (2024)
by: Bhatt, Gantavya, et al.
Published: (2024)
Competition-Level Problems are Effective LLM Evaluators
by: Huang, Yiming, et al.
Published: (2023)
by: Huang, Yiming, et al.
Published: (2023)
Cost-Effective Hallucination Detection for LLMs
by: Valentin, Simon, et al.
Published: (2024)
by: Valentin, Simon, et al.
Published: (2024)
MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards
by: Shen, Zhiyu, et al.
Published: (2026)
by: Shen, Zhiyu, et al.
Published: (2026)
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
by: Zhang, Xichen, et al.
Published: (2026)
by: Zhang, Xichen, et al.
Published: (2026)
Similar Items
-
Temperature-Centric Investigation of Speculative Decoding with Knowledge Distillation
by: Ouyang, Siru, et al.
Published: (2024) -
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
by: Wang, Yiping, et al.
Published: (2024) -
Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
by: Zhang, Shenao, et al.
Published: (2024) -
Rethinking Data Synthesis: A Teacher Model Training Recipe with Interpretation
by: Chen, Yifang, et al.
Published: (2024) -
Multi-LoRA Composition for Image Generation
by: Zhong, Ming, et al.
Published: (2024)