Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Renjun, Cheng, Yi, Meng, Libin, Xia, Jiaxin, Zong, Yi, Shi, Xing, Lin, Wei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Investigating Non-Transitivity in LLM-as-a-Judge
por: Xu, Yi, et al.
Publicado: (2025)
por: Xu, Yi, et al.
Publicado: (2025)
Large Language Models Could Be Rote Learners
por: Xu, Yuyang, et al.
Publicado: (2025)
por: Xu, Yuyang, et al.
Publicado: (2025)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
por: Hong, Fenglu, et al.
Publicado: (2025)
por: Hong, Fenglu, et al.
Publicado: (2025)
Arithmetic Feature Interaction Is Necessary for Deep Tabular Learning
por: Cheng, Yi, et al.
Publicado: (2024)
por: Cheng, Yi, et al.
Publicado: (2024)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
por: Liu, Yixin, et al.
Publicado: (2026)
por: Liu, Yixin, et al.
Publicado: (2026)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
por: Li, Zhuochun, et al.
Publicado: (2026)
por: Li, Zhuochun, et al.
Publicado: (2026)
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
por: Xu, Yuyang, et al.
Publicado: (2025)
por: Xu, Yuyang, et al.
Publicado: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
por: Yang, Jinming, et al.
Publicado: (2026)
por: Yang, Jinming, et al.
Publicado: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
por: Tan, Sijun, et al.
Publicado: (2024)
por: Tan, Sijun, et al.
Publicado: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
por: Shi, Lin, et al.
Publicado: (2024)
por: Shi, Lin, et al.
Publicado: (2024)
JTCSE: Joint Tensor-Modulus Constraints and Cross-Attention for Unsupervised Contrastive Learning of Sentence Embeddings
por: Zong, Tianyu, et al.
Publicado: (2025)
por: Zong, Tianyu, et al.
Publicado: (2025)
On Designing Effective RL Reward at Training Time for LLM Reasoning
por: Gao, Jiaxuan, et al.
Publicado: (2024)
por: Gao, Jiaxuan, et al.
Publicado: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
por: Xu, Austin, et al.
Publicado: (2025)
por: Xu, Austin, et al.
Publicado: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
por: Hong, Yihan, et al.
Publicado: (2026)
por: Hong, Yihan, et al.
Publicado: (2026)
Training-free LLM Merging for Multi-task Learning
por: Fu, Zichuan, et al.
Publicado: (2025)
por: Fu, Zichuan, et al.
Publicado: (2025)
Preparing Lessons for Progressive Training on Language Models
por: Pan, Yu, et al.
Publicado: (2024)
por: Pan, Yu, et al.
Publicado: (2024)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
por: Zhang, Zhenru, et al.
Publicado: (2025)
por: Zhang, Zhenru, et al.
Publicado: (2025)
Enhancing Hepatopathy Clinical Trial Efficiency: A Secure, Large Language Model-Powered Pre-Screening Pipeline
por: Gui, Xiongbin, et al.
Publicado: (2025)
por: Gui, Xiongbin, et al.
Publicado: (2025)
CultureLLM: Incorporating Cultural Differences into Large Language Models
por: Li, Cheng, et al.
Publicado: (2024)
por: Li, Cheng, et al.
Publicado: (2024)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
por: Hu, Xing, et al.
Publicado: (2024)
por: Hu, Xing, et al.
Publicado: (2024)
JuStRank: Benchmarking LLM Judges for System Ranking
por: Gera, Ariel, et al.
Publicado: (2024)
por: Gera, Ariel, et al.
Publicado: (2024)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
por: Salinas, David, et al.
Publicado: (2025)
por: Salinas, David, et al.
Publicado: (2025)
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
por: You, Haoran, et al.
Publicado: (2024)
por: You, Haoran, et al.
Publicado: (2024)
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
por: Huang, Wei, et al.
Publicado: (2026)
por: Huang, Wei, et al.
Publicado: (2026)
Schema Lineage Extraction at Scale: Multilingual Pipelines, Composite Evaluation, and Language-Model Benchmarks
por: Yin, Jiaqi, et al.
Publicado: (2025)
por: Yin, Jiaqi, et al.
Publicado: (2025)
Towards Best Practices for Open Datasets for LLM Training
por: Baack, Stefan, et al.
Publicado: (2025)
por: Baack, Stefan, et al.
Publicado: (2025)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
por: Jing, Yi, et al.
Publicado: (2026)
por: Jing, Yi, et al.
Publicado: (2026)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
por: Yehudai, Asaf, et al.
Publicado: (2025)
por: Yehudai, Asaf, et al.
Publicado: (2025)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
por: Gupta, Manan, et al.
Publicado: (2026)
por: Gupta, Manan, et al.
Publicado: (2026)
Learning Dynamics of LLM Finetuning
por: Ren, Yi, et al.
Publicado: (2024)
por: Ren, Yi, et al.
Publicado: (2024)
BAGEN: Are LLM Agents Budget-Aware?
por: Lin, Yuxiang, et al.
Publicado: (2026)
por: Lin, Yuxiang, et al.
Publicado: (2026)
Ask Again, Then Fail: Large Language Models' Vacillations in Judgment
por: Xie, Qiming, et al.
Publicado: (2023)
por: Xie, Qiming, et al.
Publicado: (2023)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
por: Bouchard, Dylan, et al.
Publicado: (2025)
por: Bouchard, Dylan, et al.
Publicado: (2025)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
por: Whitehouse, Chenxi, et al.
Publicado: (2025)
por: Whitehouse, Chenxi, et al.
Publicado: (2025)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
por: Landesberg, Eddie
Publicado: (2026)
por: Landesberg, Eddie
Publicado: (2026)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
por: Xu, Ran, et al.
Publicado: (2025)
por: Xu, Ran, et al.
Publicado: (2025)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
por: Zhang, Wenbo, et al.
Publicado: (2026)
por: Zhang, Wenbo, et al.
Publicado: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
por: Alam, Firoj, et al.
Publicado: (2026)
por: Alam, Firoj, et al.
Publicado: (2026)
Open or Closed LLM for Lesser-Resourced Languages? Lessons from Greek
por: Pavlopoulos, John, et al.
Publicado: (2025)
por: Pavlopoulos, John, et al.
Publicado: (2025)
Ejemplares similares
-
Investigating Non-Transitivity in LLM-as-a-Judge
por: Xu, Yi, et al.
Publicado: (2025) -
Large Language Models Could Be Rote Learners
por: Xu, Yuyang, et al.
Publicado: (2025) -
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
por: Hong, Fenglu, et al.
Publicado: (2025) -
Arithmetic Feature Interaction Is Necessary for Deep Tabular Learning
por: Cheng, Yi, et al.
Publicado: (2024) -
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
por: Liu, Yixin, et al.
Publicado: (2026)