AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling
Fuente:
arXiv
Salvato in:
| Autori principali: | Miao, Yongliang, Liang, Yangyang, Du, Mengnan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models
di: Liu, Weiqi, et al.
Pubblicazione: (2026)
di: Liu, Weiqi, et al.
Pubblicazione: (2026)
Debate Helps Weak Judges Reward Stronger Models
di: Elasky, Ethan, et al.
Pubblicazione: (2026)
di: Elasky, Ethan, et al.
Pubblicazione: (2026)
AutoJudge: Judge Decoding Without Manual Annotation
di: Garipov, Roman, et al.
Pubblicazione: (2025)
di: Garipov, Roman, et al.
Pubblicazione: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
di: Duo, Jiangshan, et al.
Pubblicazione: (2026)
di: Duo, Jiangshan, et al.
Pubblicazione: (2026)
Judge Circuits
di: Feldhus, Nils, et al.
Pubblicazione: (2026)
di: Feldhus, Nils, et al.
Pubblicazione: (2026)
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
di: Ye, Ziyi, et al.
Pubblicazione: (2024)
di: Ye, Ziyi, et al.
Pubblicazione: (2024)
Quantitative LLM Judges
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment
di: Bachmann, Gregor, et al.
Pubblicazione: (2025)
di: Bachmann, Gregor, et al.
Pubblicazione: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
di: Li, Zhuochun, et al.
Pubblicazione: (2026)
di: Li, Zhuochun, et al.
Pubblicazione: (2026)
One Token to Fool LLM-as-a-Judge
di: Zhao, Yulai, et al.
Pubblicazione: (2025)
di: Zhao, Yulai, et al.
Pubblicazione: (2025)
AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining
di: Dong, Hongyuan, et al.
Pubblicazione: (2025)
di: Dong, Hongyuan, et al.
Pubblicazione: (2025)
JAF: Judge Agent Forest
di: Garg, Sahil, et al.
Pubblicazione: (2026)
di: Garg, Sahil, et al.
Pubblicazione: (2026)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Learning to Judge: LLMs Designing and Applying Evaluation Rubrics
di: Siro, Clemencia, et al.
Pubblicazione: (2026)
di: Siro, Clemencia, et al.
Pubblicazione: (2026)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
di: Zhang, Wenbo, et al.
Pubblicazione: (2026)
di: Zhang, Wenbo, et al.
Pubblicazione: (2026)
AdaSplash: Adaptive Sparse Flash Attention
di: Gonçalves, Nuno, et al.
Pubblicazione: (2025)
di: Gonçalves, Nuno, et al.
Pubblicazione: (2025)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
di: Li, Tung-Ling, et al.
Pubblicazione: (2025)
di: Li, Tung-Ling, et al.
Pubblicazione: (2025)
Approximating Human Preferences Using a Multi-Judge Learned System
di: Sprejer, Eitán, et al.
Pubblicazione: (2025)
di: Sprejer, Eitán, et al.
Pubblicazione: (2025)
Reinforcement Learning-based Knowledge Distillation with LLM-as-a-Judge
di: Shen, Yiyang, et al.
Pubblicazione: (2026)
di: Shen, Yiyang, et al.
Pubblicazione: (2026)
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
di: Sutawika, Lintang, et al.
Pubblicazione: (2026)
di: Sutawika, Lintang, et al.
Pubblicazione: (2026)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
The Judge Variable: Challenging Judge-Agnostic Legal Judgment Prediction
di: Zambrano, Guillaume
Pubblicazione: (2025)
di: Zambrano, Guillaume
Pubblicazione: (2025)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
di: Ballon, Marthe, et al.
Pubblicazione: (2026)
di: Ballon, Marthe, et al.
Pubblicazione: (2026)
SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection
di: Zhang, Huopu, et al.
Pubblicazione: (2025)
di: Zhang, Huopu, et al.
Pubblicazione: (2025)
CodeJudge: Evaluating Code Generation with Large Language Models
di: Tong, Weixi, et al.
Pubblicazione: (2024)
di: Tong, Weixi, et al.
Pubblicazione: (2024)
How to Correctly Report LLM-as-a-Judge Evaluations
di: Lee, Chungpa, et al.
Pubblicazione: (2025)
di: Lee, Chungpa, et al.
Pubblicazione: (2025)
Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
di: Wei, Zhipeng, et al.
Pubblicazione: (2024)
di: Wei, Zhipeng, et al.
Pubblicazione: (2024)
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
di: Wang, Zhilin, et al.
Pubblicazione: (2025)
di: Wang, Zhilin, et al.
Pubblicazione: (2025)
Comparative Analysis of Demonstration Selection Algorithms for LLM In-Context Learning
di: Shu, Dong, et al.
Pubblicazione: (2024)
di: Shu, Dong, et al.
Pubblicazione: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
The Perfect Blend: Redefining RLHF with Mixture of Judges
di: Xu, Tengyu, et al.
Pubblicazione: (2024)
di: Xu, Tengyu, et al.
Pubblicazione: (2024)
Investigating Non-Transitivity in LLM-as-a-Judge
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
AdaLomo: Low-memory Optimization with Adaptive Learning Rate
di: Lv, Kai, et al.
Pubblicazione: (2023)
di: Lv, Kai, et al.
Pubblicazione: (2023)
LASeR: Learning to Adaptively Select Reward Models with Multi-Armed Bandits
di: Nguyen, Duy, et al.
Pubblicazione: (2024)
di: Nguyen, Duy, et al.
Pubblicazione: (2024)
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
di: Hu, Renjun, et al.
Pubblicazione: (2025)
di: Hu, Renjun, et al.
Pubblicazione: (2025)
Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
Silence the Judge: Reinforcement Learning with Self-Verifier via Latent Geometric Clustering
di: Zhang, Nonghai, et al.
Pubblicazione: (2026)
di: Zhang, Nonghai, et al.
Pubblicazione: (2026)
Documenti analoghi
-
NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models
di: Liu, Weiqi, et al.
Pubblicazione: (2026) -
Debate Helps Weak Judges Reward Stronger Models
di: Elasky, Ethan, et al.
Pubblicazione: (2026) -
AutoJudge: Judge Decoding Without Manual Annotation
di: Garipov, Roman, et al.
Pubblicazione: (2025) -
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
di: Duo, Jiangshan, et al.
Pubblicazione: (2026) -
Judge Circuits
di: Feldhus, Nils, et al.
Pubblicazione: (2026)