WildReward: Learning Reward Models from In-the-Wild Human Interactions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Peng, Hao, Qi, Yunjia, Wang, Xiaozhi, Yao, Zijun, Hou, Lei, Li, Juanzi |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
par: Peng, Hao, et autres
Publié: (2025)
par: Peng, Hao, et autres
Publié: (2025)
StoryAlign: Evaluating and Training Reward Models for Story Generation
par: Xia, Haotian, et autres
Publié: (2026)
par: Xia, Haotian, et autres
Publié: (2026)
VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
par: Peng, Hao, et autres
Publié: (2025)
par: Peng, Hao, et autres
Publié: (2025)
Constraint Back-translation Improves Complex Instruction Following of Large Language Models
par: Qi, Yunjia, et autres
Publié: (2024)
par: Qi, Yunjia, et autres
Publié: (2024)
StoryWriter: A Multi-Agent Framework for Long Story Generation
par: Xia, Haotian, et autres
Publié: (2025)
par: Xia, Haotian, et autres
Publié: (2025)
Auxiliary Metrics Help Decoding Skill Neurons in the Wild
par: Zhao, Yixiu, et autres
Publié: (2025)
par: Zhao, Yixiu, et autres
Publié: (2025)
AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
par: Qi, Yunjia, et autres
Publié: (2025)
par: Qi, Yunjia, et autres
Publié: (2025)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
par: Chen, Jianhui, et autres
Publié: (2024)
par: Chen, Jianhui, et autres
Publié: (2024)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
par: Jing, Yi, et autres
Publié: (2026)
par: Jing, Yi, et autres
Publié: (2026)
Pre-training Distillation for Large Language Models: A Design Space Exploration
par: Peng, Hao, et autres
Publié: (2024)
par: Peng, Hao, et autres
Publié: (2024)
ADELIE: Aligning Large Language Models on Information Extraction
par: Qi, Yunjia, et autres
Publié: (2024)
par: Qi, Yunjia, et autres
Publié: (2024)
Rewarding Creativity: A Human-Aligned Generative Reward Model for Reinforcement Learning in Storytelling
par: Li, Zhaoyan, et autres
Publié: (2026)
par: Li, Zhaoyan, et autres
Publié: (2026)
Event-level Knowledge Editing
par: Peng, Hao, et autres
Publié: (2024)
par: Peng, Hao, et autres
Publié: (2024)
ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time
par: Tu, Shangqing, et autres
Publié: (2023)
par: Tu, Shangqing, et autres
Publié: (2023)
MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark
par: Ma, Shengkun, et autres
Publié: (2025)
par: Ma, Shengkun, et autres
Publié: (2025)
RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
par: Liu, Yantao, et autres
Publié: (2024)
par: Liu, Yantao, et autres
Publié: (2024)
WildSci: Advancing Scientific Reasoning from In-the-Wild Literature
par: Liu, Tengxiao, et autres
Publié: (2026)
par: Liu, Tengxiao, et autres
Publié: (2026)
Inference-Time Scaling for Generalist Reward Modeling
par: Liu, Zijun, et autres
Publié: (2025)
par: Liu, Zijun, et autres
Publié: (2025)
TacoERE: Cluster-aware Compression for Event Relation Extraction
par: Guan, Yong, et autres
Publié: (2024)
par: Guan, Yong, et autres
Publié: (2024)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
par: Lu, Yujie, et autres
Publié: (2024)
par: Lu, Yujie, et autres
Publié: (2024)
MAVEN-Fact: A Large-scale Event Factuality Detection Dataset
par: Li, Chunyang, et autres
Publié: (2024)
par: Li, Chunyang, et autres
Publié: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
par: Fu, Jiayi, et autres
Publié: (2025)
par: Fu, Jiayi, et autres
Publié: (2025)
A Cause-Effect Look at Alleviating Hallucination of Knowledge-grounded Dialogue Generation
par: Yu, Jifan, et autres
Publié: (2024)
par: Yu, Jifan, et autres
Publié: (2024)
Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards
par: Shen, Wei, et autres
Publié: (2024)
par: Shen, Wei, et autres
Publié: (2024)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
par: Tu, Shangqing, et autres
Publié: (2024)
par: Tu, Shangqing, et autres
Publié: (2024)
WildIFEval: Instruction Following in the Wild
par: Lior, Gili, et autres
Publié: (2025)
par: Lior, Gili, et autres
Publié: (2025)
Reverse That Number! Decoding Order Matters in Arithmetic Learning
par: Zhang-Li, Daniel, et autres
Publié: (2024)
par: Zhang-Li, Daniel, et autres
Publié: (2024)
Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
par: Elle
Publié: (2025)
par: Elle
Publié: (2025)
MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models
par: Tang, Zecheng, et autres
Publié: (2026)
par: Tang, Zecheng, et autres
Publié: (2026)
GRAM: A Generative Foundation Reward Model for Reward Generalization
par: Wang, Chenglong, et autres
Publié: (2025)
par: Wang, Chenglong, et autres
Publié: (2025)
Exploring Reasoning Reward Model for Agents
par: Fan, Kaixuan, et autres
Publié: (2026)
par: Fan, Kaixuan, et autres
Publié: (2026)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
par: Liu, Chris Yuhao, et autres
Publié: (2024)
par: Liu, Chris Yuhao, et autres
Publié: (2024)
Process Reward Models That Think
par: Khalifa, Muhammad, et autres
Publié: (2025)
par: Khalifa, Muhammad, et autres
Publié: (2025)
Fine-Tuning Language Models with Reward Learning on Policy
par: Lang, Hao, et autres
Publié: (2024)
par: Lang, Hao, et autres
Publié: (2024)
Self-Evolved Reward Learning for LLMs
par: Huang, Chenghua, et autres
Publié: (2024)
par: Huang, Chenghua, et autres
Publié: (2024)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
par: Lin, Bill Yuchen, et autres
Publié: (2024)
par: Lin, Bill Yuchen, et autres
Publié: (2024)
Reward Models Identify Consistency, Not Causality
par: Xu, Yuhui, et autres
Publié: (2025)
par: Xu, Yuhui, et autres
Publié: (2025)
Reward Is Enough: LLMs Are In-Context Reinforcement Learners
par: Song, Kefan, et autres
Publié: (2025)
par: Song, Kefan, et autres
Publié: (2025)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
par: Ackermann, Johannes, et autres
Publié: (2026)
par: Ackermann, Johannes, et autres
Publié: (2026)
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
par: Cao, Meng, et autres
Publié: (2024)
par: Cao, Meng, et autres
Publié: (2024)
Documents similaires
-
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
par: Peng, Hao, et autres
Publié: (2025) -
StoryAlign: Evaluating and Training Reward Models for Story Generation
par: Xia, Haotian, et autres
Publié: (2026) -
VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
par: Peng, Hao, et autres
Publié: (2025) -
Constraint Back-translation Improves Complex Instruction Following of Large Language Models
par: Qi, Yunjia, et autres
Publié: (2024) -
StoryWriter: A Multi-Agent Framework for Long Story Generation
par: Xia, Haotian, et autres
Publié: (2025)