Exploring Domain Robust Lightweight Reward Models based on Router Mechanism
Fuente:
arXiv
Salvato in:
| Autori principali: | Namgoong, Hyuk, Jung, Jeesu, Jung, Sangkeun, Roh, Yoonhyung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
Reasoning Steps as Curriculum: Using Depth of Thought as a Difficulty Signal for Tuning LLMs
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
Aggregated Knowledge Model: Enhancing Domain-Specific QA with Fine-Tuned and Retrieval-Augmented Generation Models
di: Liu, Fengchen, et al.
Pubblicazione: (2024)
di: Liu, Fengchen, et al.
Pubblicazione: (2024)
RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models
di: Chen, Shuhao, et al.
Pubblicazione: (2024)
di: Chen, Shuhao, et al.
Pubblicazione: (2024)
Evaluating Large Language Models on the 2026 Korean CSAT Mathematics Exam: Measuring Mathematical Ability in a Zero-Data-Leakage Setting
di: Pyeon, Goun, et al.
Pubblicazione: (2025)
di: Pyeon, Goun, et al.
Pubblicazione: (2025)
OrcaRouter: A Production-Oriented LLM Router with Hybrid Offline-Online Learning
di: Bao, Zhenghua, et al.
Pubblicazione: (2026)
di: Bao, Zhenghua, et al.
Pubblicazione: (2026)
On the Robustness of Reward Models for Language Model Alignment
di: Hong, Jiwoo, et al.
Pubblicazione: (2025)
di: Hong, Jiwoo, et al.
Pubblicazione: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
di: Kim, Sunghwan, et al.
Pubblicazione: (2024)
di: Kim, Sunghwan, et al.
Pubblicazione: (2024)
LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis
di: Heo, Inbum, et al.
Pubblicazione: (2026)
di: Heo, Inbum, et al.
Pubblicazione: (2026)
Employing Layerwised Unsupervised Learning to Lessen Data and Loss Requirements in Forward-Forward Algorithms
di: Hwang, Taewook, et al.
Pubblicazione: (2024)
di: Hwang, Taewook, et al.
Pubblicazione: (2024)
VL-RouterBench: A Benchmark for Vision-Language Model Routing
di: Huang, Zhehao, et al.
Pubblicazione: (2025)
di: Huang, Zhehao, et al.
Pubblicazione: (2025)
R3: Robust Rubric-Agnostic Reward Models
di: Anugraha, David, et al.
Pubblicazione: (2025)
di: Anugraha, David, et al.
Pubblicazione: (2025)
Reward-Robust RLHF in LLMs
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
Exploring RL-based LLM Training for Formal Language Tasks with Programmed Rewards
di: Padula, Alexander G., et al.
Pubblicazione: (2024)
di: Padula, Alexander G., et al.
Pubblicazione: (2024)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
di: Gunjal, Anisha, et al.
Pubblicazione: (2025)
di: Gunjal, Anisha, et al.
Pubblicazione: (2025)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
di: Ma, Wenhan, et al.
Pubblicazione: (2025)
di: Ma, Wenhan, et al.
Pubblicazione: (2025)
Phasor Memory Networks: Stable Backpropagation Through Time for Scalable Explicit Memory
di: Goo, Sungwoo, et al.
Pubblicazione: (2026)
di: Goo, Sungwoo, et al.
Pubblicazione: (2026)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
di: Xie, Yanyue, et al.
Pubblicazione: (2024)
di: Xie, Yanyue, et al.
Pubblicazione: (2024)
RewardAnything: Generalizable Principle-Following Reward Models
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
MathBridge: A Large Corpus Dataset for Translating Spoken Mathematical Expressions into $LaTeX$ Formulas for Improved Readability
di: Jung, Kyudan, et al.
Pubblicazione: (2024)
di: Jung, Kyudan, et al.
Pubblicazione: (2024)
Guidance-Based Prompt Data Augmentation in Specialized Domains for Named Entity Recognition
di: Kang, Hyeonseok, et al.
Pubblicazione: (2024)
di: Kang, Hyeonseok, et al.
Pubblicazione: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
di: Ono, Shinnosuke, et al.
Pubblicazione: (2026)
di: Ono, Shinnosuke, et al.
Pubblicazione: (2026)
How Much is Too Much? Exploring LoRA Rank Trade-offs for Retaining Knowledge and Domain Robustness
di: Rathore, Darshita, et al.
Pubblicazione: (2025)
di: Rathore, Darshita, et al.
Pubblicazione: (2025)
M-RewardBench: Evaluating Reward Models in Multilingual Settings
di: Gureja, Srishti, et al.
Pubblicazione: (2024)
di: Gureja, Srishti, et al.
Pubblicazione: (2024)
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
di: Kim, Sunghwan, et al.
Pubblicazione: (2025)
di: Kim, Sunghwan, et al.
Pubblicazione: (2025)
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
di: Xu, Xin, et al.
Pubblicazione: (2026)
di: Xu, Xin, et al.
Pubblicazione: (2026)
xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
di: Qian, Cheng, et al.
Pubblicazione: (2025)
di: Qian, Cheng, et al.
Pubblicazione: (2025)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
di: Yang, Daniel, et al.
Pubblicazione: (2026)
di: Yang, Daniel, et al.
Pubblicazione: (2026)
Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined Data
di: Ling, Zhenqing, et al.
Pubblicazione: (2025)
di: Ling, Zhenqing, et al.
Pubblicazione: (2025)
Process Reward Models That Think
di: Khalifa, Muhammad, et al.
Pubblicazione: (2025)
di: Khalifa, Muhammad, et al.
Pubblicazione: (2025)
Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning
di: Zhang, Haozhen, et al.
Pubblicazione: (2025)
di: Zhang, Haozhen, et al.
Pubblicazione: (2025)
Robust Probabilistic Model Checking with Continuous Reward Domains
di: Ji, Xiaotong, et al.
Pubblicazione: (2025)
di: Ji, Xiaotong, et al.
Pubblicazione: (2025)
Binary Classifier Optimization for Large Language Model Alignment
di: Jung, Seungjae, et al.
Pubblicazione: (2024)
di: Jung, Seungjae, et al.
Pubblicazione: (2024)
How to Evaluate Reward Models for RLHF
di: Frick, Evan, et al.
Pubblicazione: (2024)
di: Frick, Evan, et al.
Pubblicazione: (2024)
Reward Model Overoptimisation in Iterated RLHF
di: Wolf, Lorenz, et al.
Pubblicazione: (2025)
di: Wolf, Lorenz, et al.
Pubblicazione: (2025)
Reward Models Identify Consistency, Not Causality
di: Xu, Yuhui, et al.
Pubblicazione: (2025)
di: Xu, Yuhui, et al.
Pubblicazione: (2025)
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
di: Wang, Zhilin, et al.
Pubblicazione: (2025)
di: Wang, Zhilin, et al.
Pubblicazione: (2025)
Explicit Diversity Conditions for Effective Question Answer Generation with Large Language Models
di: Yadav, Vikas, et al.
Pubblicazione: (2024)
di: Yadav, Vikas, et al.
Pubblicazione: (2024)
Toward Super Agent System with Hybrid AI Routers
di: Yao, Yuhang, et al.
Pubblicazione: (2025)
di: Yao, Yuhang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction
di: Jung, Jeesu, et al.
Pubblicazione: (2025) -
Reasoning Steps as Curriculum: Using Depth of Thought as a Difficulty Signal for Tuning LLMs
di: Jung, Jeesu, et al.
Pubblicazione: (2025) -
Aggregated Knowledge Model: Enhancing Domain-Specific QA with Fine-Tuned and Retrieval-Augmented Generation Models
di: Liu, Fengchen, et al.
Pubblicazione: (2024) -
RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models
di: Chen, Shuhao, et al.
Pubblicazione: (2024) -
Evaluating Large Language Models on the 2026 Korean CSAT Mathematics Exam: Measuring Mathematical Ability in a Zero-Data-Leakage Setting
di: Pyeon, Goun, et al.
Pubblicazione: (2025)