Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Renjun, Cheng, Yi, Meng, Libin, Xia, Jiaxin, Zong, Yi, Shi, Xing, Lin, Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Investigating Non-Transitivity in LLM-as-a-Judge
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
Large Language Models Could Be Rote Learners
di: Xu, Yuyang, et al.
Pubblicazione: (2025)
di: Xu, Yuyang, et al.
Pubblicazione: (2025)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
di: Hong, Fenglu, et al.
Pubblicazione: (2025)
di: Hong, Fenglu, et al.
Pubblicazione: (2025)
Arithmetic Feature Interaction Is Necessary for Deep Tabular Learning
di: Cheng, Yi, et al.
Pubblicazione: (2024)
di: Cheng, Yi, et al.
Pubblicazione: (2024)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
di: Liu, Yixin, et al.
Pubblicazione: (2026)
di: Liu, Yixin, et al.
Pubblicazione: (2026)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
di: Li, Zhuochun, et al.
Pubblicazione: (2026)
di: Li, Zhuochun, et al.
Pubblicazione: (2026)
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
di: Xu, Yuyang, et al.
Pubblicazione: (2025)
di: Xu, Yuyang, et al.
Pubblicazione: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
di: Yang, Jinming, et al.
Pubblicazione: (2026)
di: Yang, Jinming, et al.
Pubblicazione: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
di: Shi, Lin, et al.
Pubblicazione: (2024)
di: Shi, Lin, et al.
Pubblicazione: (2024)
JTCSE: Joint Tensor-Modulus Constraints and Cross-Attention for Unsupervised Contrastive Learning of Sentence Embeddings
di: Zong, Tianyu, et al.
Pubblicazione: (2025)
di: Zong, Tianyu, et al.
Pubblicazione: (2025)
On Designing Effective RL Reward at Training Time for LLM Reasoning
di: Gao, Jiaxuan, et al.
Pubblicazione: (2024)
di: Gao, Jiaxuan, et al.
Pubblicazione: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
di: Hong, Yihan, et al.
Pubblicazione: (2026)
di: Hong, Yihan, et al.
Pubblicazione: (2026)
Training-free LLM Merging for Multi-task Learning
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
Preparing Lessons for Progressive Training on Language Models
di: Pan, Yu, et al.
Pubblicazione: (2024)
di: Pan, Yu, et al.
Pubblicazione: (2024)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
di: Zhang, Zhenru, et al.
Pubblicazione: (2025)
di: Zhang, Zhenru, et al.
Pubblicazione: (2025)
Enhancing Hepatopathy Clinical Trial Efficiency: A Secure, Large Language Model-Powered Pre-Screening Pipeline
di: Gui, Xiongbin, et al.
Pubblicazione: (2025)
di: Gui, Xiongbin, et al.
Pubblicazione: (2025)
CultureLLM: Incorporating Cultural Differences into Large Language Models
di: Li, Cheng, et al.
Pubblicazione: (2024)
di: Li, Cheng, et al.
Pubblicazione: (2024)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
di: Hu, Xing, et al.
Pubblicazione: (2024)
di: Hu, Xing, et al.
Pubblicazione: (2024)
JuStRank: Benchmarking LLM Judges for System Ranking
di: Gera, Ariel, et al.
Pubblicazione: (2024)
di: Gera, Ariel, et al.
Pubblicazione: (2024)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
di: Salinas, David, et al.
Pubblicazione: (2025)
di: Salinas, David, et al.
Pubblicazione: (2025)
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
di: You, Haoran, et al.
Pubblicazione: (2024)
di: You, Haoran, et al.
Pubblicazione: (2024)
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
di: Huang, Wei, et al.
Pubblicazione: (2026)
di: Huang, Wei, et al.
Pubblicazione: (2026)
Schema Lineage Extraction at Scale: Multilingual Pipelines, Composite Evaluation, and Language-Model Benchmarks
di: Yin, Jiaqi, et al.
Pubblicazione: (2025)
di: Yin, Jiaqi, et al.
Pubblicazione: (2025)
Towards Best Practices for Open Datasets for LLM Training
di: Baack, Stefan, et al.
Pubblicazione: (2025)
di: Baack, Stefan, et al.
Pubblicazione: (2025)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
di: Jing, Yi, et al.
Pubblicazione: (2026)
di: Jing, Yi, et al.
Pubblicazione: (2026)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
di: Gupta, Manan, et al.
Pubblicazione: (2026)
di: Gupta, Manan, et al.
Pubblicazione: (2026)
Learning Dynamics of LLM Finetuning
di: Ren, Yi, et al.
Pubblicazione: (2024)
di: Ren, Yi, et al.
Pubblicazione: (2024)
BAGEN: Are LLM Agents Budget-Aware?
di: Lin, Yuxiang, et al.
Pubblicazione: (2026)
di: Lin, Yuxiang, et al.
Pubblicazione: (2026)
Ask Again, Then Fail: Large Language Models' Vacillations in Judgment
di: Xie, Qiming, et al.
Pubblicazione: (2023)
di: Xie, Qiming, et al.
Pubblicazione: (2023)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
di: Bouchard, Dylan, et al.
Pubblicazione: (2025)
di: Bouchard, Dylan, et al.
Pubblicazione: (2025)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2025)
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2025)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
di: Landesberg, Eddie
Pubblicazione: (2026)
di: Landesberg, Eddie
Pubblicazione: (2026)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
di: Xu, Ran, et al.
Pubblicazione: (2025)
di: Xu, Ran, et al.
Pubblicazione: (2025)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
di: Zhang, Wenbo, et al.
Pubblicazione: (2026)
di: Zhang, Wenbo, et al.
Pubblicazione: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
di: Alam, Firoj, et al.
Pubblicazione: (2026)
di: Alam, Firoj, et al.
Pubblicazione: (2026)
Open or Closed LLM for Lesser-Resourced Languages? Lessons from Greek
di: Pavlopoulos, John, et al.
Pubblicazione: (2025)
di: Pavlopoulos, John, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Investigating Non-Transitivity in LLM-as-a-Judge
di: Xu, Yi, et al.
Pubblicazione: (2025) -
Large Language Models Could Be Rote Learners
di: Xu, Yuyang, et al.
Pubblicazione: (2025) -
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
di: Hong, Fenglu, et al.
Pubblicazione: (2025) -
Arithmetic Feature Interaction Is Necessary for Deep Tabular Learning
di: Cheng, Yi, et al.
Pubblicazione: (2024) -
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
di: Liu, Yixin, et al.
Pubblicazione: (2026)