MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Yutong, Ji, Pengliang, Yang, Chaoqun, Li, Kaixin, Hu, Ming, Li, Jiaoyang, Sartoretti, Guillaume |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
por: Li, Xiaochuan, et al.
Publicado: (2025)
por: Li, Xiaochuan, et al.
Publicado: (2025)
Beyond Policy Optimization: A Data Curation Flywheel for Sparse-Reward Long-Horizon Planning
por: Wang, Yutong, et al.
Publicado: (2025)
por: Wang, Yutong, et al.
Publicado: (2025)
Deploying Ten Thousand Robots: Scalable Imitation Learning for Lifelong Multi-Agent Path Finding
por: Jiang, He, et al.
Publicado: (2024)
por: Jiang, He, et al.
Publicado: (2024)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
por: Huang, Tzu-Heng, et al.
Publicado: (2025)
por: Huang, Tzu-Heng, et al.
Publicado: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
por: Xu, Austin, et al.
Publicado: (2025)
por: Xu, Austin, et al.
Publicado: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
por: Tan, Sijun, et al.
Publicado: (2024)
por: Tan, Sijun, et al.
Publicado: (2024)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
por: Li, Zhuochun, et al.
Publicado: (2026)
por: Li, Zhuochun, et al.
Publicado: (2026)
Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
por: Halder, Indranil, et al.
Publicado: (2025)
por: Halder, Indranil, et al.
Publicado: (2025)
Auto-Prompt Ensemble for LLM Judge
por: Li, Jiajie, et al.
Publicado: (2025)
por: Li, Jiajie, et al.
Publicado: (2025)
Becoming Experienced Judges: Selective Test-Time Learning for Evaluators
por: Jwa, Seungyeon, et al.
Publicado: (2025)
por: Jwa, Seungyeon, et al.
Publicado: (2025)
To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay
por: Dey, Soumik, et al.
Publicado: (2025)
por: Dey, Soumik, et al.
Publicado: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
por: Yang, Jinming, et al.
Publicado: (2026)
por: Yang, Jinming, et al.
Publicado: (2026)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
por: Duo, Jiangshan, et al.
Publicado: (2026)
por: Duo, Jiangshan, et al.
Publicado: (2026)
S*: Test Time Scaling for Code Generation
por: Li, Dacheng, et al.
Publicado: (2025)
por: Li, Dacheng, et al.
Publicado: (2025)
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking
por: Liu, Zijun, et al.
Publicado: (2024)
por: Liu, Zijun, et al.
Publicado: (2024)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
por: Radharapu, Bhaktipriya, et al.
Publicado: (2025)
por: Radharapu, Bhaktipriya, et al.
Publicado: (2025)
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
por: Zhu, Xiao, et al.
Publicado: (2026)
por: Zhu, Xiao, et al.
Publicado: (2026)
Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge
por: Fabbri, Francesco, et al.
Publicado: (2025)
por: Fabbri, Francesco, et al.
Publicado: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
por: Zhou, Yilun, et al.
Publicado: (2025)
por: Zhou, Yilun, et al.
Publicado: (2025)
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
por: Wagner, Nico, et al.
Publicado: (2024)
por: Wagner, Nico, et al.
Publicado: (2024)
REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge
por: Zhang, Yasi, et al.
Publicado: (2026)
por: Zhang, Yasi, et al.
Publicado: (2026)
Investigating Non-Transitivity in LLM-as-a-Judge
por: Xu, Yi, et al.
Publicado: (2025)
por: Xu, Yi, et al.
Publicado: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
por: Hong, Yihan, et al.
Publicado: (2026)
por: Hong, Yihan, et al.
Publicado: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
por: Alam, Firoj, et al.
Publicado: (2026)
por: Alam, Firoj, et al.
Publicado: (2026)
Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge
por: Dadkhahi, Hamid, et al.
Publicado: (2025)
por: Dadkhahi, Hamid, et al.
Publicado: (2025)
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
por: Feuer, Benjamin, et al.
Publicado: (2024)
por: Feuer, Benjamin, et al.
Publicado: (2024)
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
por: Hu, Renjun, et al.
Publicado: (2025)
por: Hu, Renjun, et al.
Publicado: (2025)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
por: Muhamed, Aashiq
Publicado: (2025)
por: Muhamed, Aashiq
Publicado: (2025)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
por: Whitehouse, Chenxi, et al.
Publicado: (2025)
por: Whitehouse, Chenxi, et al.
Publicado: (2025)
PHUDGE: Phi-3 as Scalable Judge
por: Deshwal, Mahesh, et al.
Publicado: (2024)
por: Deshwal, Mahesh, et al.
Publicado: (2024)
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance
por: Agnihotri, Rudransh, et al.
Publicado: (2025)
por: Agnihotri, Rudransh, et al.
Publicado: (2025)
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
por: Shen, William F., et al.
Publicado: (2026)
por: Shen, William F., et al.
Publicado: (2026)
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
por: Ai, Jiaxin, et al.
Publicado: (2025)
por: Ai, Jiaxin, et al.
Publicado: (2025)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
por: Collot, Stephane, et al.
Publicado: (2025)
por: Collot, Stephane, et al.
Publicado: (2025)
DAJ: Data-Reweighted LLM Judge for Test-Time Scaling in Code Generation
por: Qin, Peijia, et al.
Publicado: (2026)
por: Qin, Peijia, et al.
Publicado: (2026)
JuStRank: Benchmarking LLM Judges for System Ranking
por: Gera, Ariel, et al.
Publicado: (2024)
por: Gera, Ariel, et al.
Publicado: (2024)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
por: Salinas, David, et al.
Publicado: (2025)
por: Salinas, David, et al.
Publicado: (2025)
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
por: Yang, Xianglin, et al.
Publicado: (2026)
por: Yang, Xianglin, et al.
Publicado: (2026)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
por: Xu, Ran, et al.
Publicado: (2025)
por: Xu, Ran, et al.
Publicado: (2025)
Ejemplares similares
-
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
por: Li, Xiaochuan, et al.
Publicado: (2025) -
Beyond Policy Optimization: A Data Curation Flywheel for Sparse-Reward Long-Horizon Planning
por: Wang, Yutong, et al.
Publicado: (2025) -
Deploying Ten Thousand Robots: Scalable Imitation Learning for Lifelong Multi-Agent Path Finding
por: Jiang, He, et al.
Publicado: (2024) -
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
por: Huang, Tzu-Heng, et al.
Publicado: (2025) -
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
por: Xu, Austin, et al.
Publicado: (2025)