Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Junjie, Kan, Xuan, He, Zihao, Tan, Shunwen, Pan, Bo, Zhang, Kaitai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Bi-Level Prompt Optimization for Multimodal LLM-as-a-Judge
di: Pan, Bo, et al.
Pubblicazione: (2026)
di: Pan, Bo, et al.
Pubblicazione: (2026)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
di: Jiang, Hongchao, et al.
Pubblicazione: (2025)
di: Jiang, Hongchao, et al.
Pubblicazione: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
di: Wan, Zhongwei, et al.
Pubblicazione: (2025)
di: Wan, Zhongwei, et al.
Pubblicazione: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
di: Yang, Bo, et al.
Pubblicazione: (2026)
di: Yang, Bo, et al.
Pubblicazione: (2026)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
di: Tang, Zhenwei, et al.
Pubblicazione: (2026)
di: Tang, Zhenwei, et al.
Pubblicazione: (2026)
Reinforcement Learning-based Knowledge Distillation with LLM-as-a-Judge
di: Shen, Yiyang, et al.
Pubblicazione: (2026)
di: Shen, Yiyang, et al.
Pubblicazione: (2026)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
di: Wei, Hui, et al.
Pubblicazione: (2024)
di: Wei, Hui, et al.
Pubblicazione: (2024)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
di: Chen, Junjie, et al.
Pubblicazione: (2026)
di: Chen, Junjie, et al.
Pubblicazione: (2026)
DELTA: Deliberative Multi-Agent Reasoning with Reinforcement Learning for Multimodal Psychological Counseling
di: Yang, Jiangnan, et al.
Pubblicazione: (2026)
di: Yang, Jiangnan, et al.
Pubblicazione: (2026)
Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
di: Le, Hieu Xuan, et al.
Pubblicazione: (2026)
di: Le, Hieu Xuan, et al.
Pubblicazione: (2026)
MR. Judge: Multimodal Reasoner as a Judge
di: Pi, Renjie, et al.
Pubblicazione: (2025)
di: Pi, Renjie, et al.
Pubblicazione: (2025)
Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges
di: Tang, Yuqi, et al.
Pubblicazione: (2025)
di: Tang, Yuqi, et al.
Pubblicazione: (2025)
Think-J: Learning to Think for Generative LLM-as-a-Judge
di: Huang, Hui, et al.
Pubblicazione: (2025)
di: Huang, Hui, et al.
Pubblicazione: (2025)
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge
di: Zhou, Karen, et al.
Pubblicazione: (2026)
di: Zhou, Karen, et al.
Pubblicazione: (2026)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
di: Xu, Ran, et al.
Pubblicazione: (2025)
di: Xu, Ran, et al.
Pubblicazione: (2025)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
di: Wang, Yidong, et al.
Pubblicazione: (2025)
di: Wang, Yidong, et al.
Pubblicazione: (2025)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
di: Bellibatlu, Rohith Reddy, et al.
Pubblicazione: (2026)
di: Bellibatlu, Rohith Reddy, et al.
Pubblicazione: (2026)
CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models
di: Liu, Wentao, et al.
Pubblicazione: (2024)
di: Liu, Wentao, et al.
Pubblicazione: (2024)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2025)
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2025)
A Survey on LLM-as-a-Judge
di: Gu, Jiawei, et al.
Pubblicazione: (2024)
di: Gu, Jiawei, et al.
Pubblicazione: (2024)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
di: Marioriyad, Arash, et al.
Pubblicazione: (2025)
di: Marioriyad, Arash, et al.
Pubblicazione: (2025)
Task Selection and Assignment for Multi-modal Multi-task Dialogue Act Classification with Non-stationary Multi-armed Bandits
di: He, Xiangheng, et al.
Pubblicazione: (2023)
di: He, Xiangheng, et al.
Pubblicazione: (2023)
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?
di: Chen, Jiamin, et al.
Pubblicazione: (2026)
di: Chen, Jiamin, et al.
Pubblicazione: (2026)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
di: Belmadani, Ikram, et al.
Pubblicazione: (2026)
di: Belmadani, Ikram, et al.
Pubblicazione: (2026)
Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards
di: Wei, Xiaolong, et al.
Pubblicazione: (2025)
di: Wei, Xiaolong, et al.
Pubblicazione: (2025)
Silence the Judge: Reinforcement Learning with Self-Verifier via Latent Geometric Clustering
di: Zhang, Nonghai, et al.
Pubblicazione: (2026)
di: Zhang, Nonghai, et al.
Pubblicazione: (2026)
Debatrix: Multi-dimensional Debate Judge with Iterative Chronological Analysis Based on LLM
di: Liang, Jingcong, et al.
Pubblicazione: (2024)
di: Liang, Jingcong, et al.
Pubblicazione: (2024)
Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
di: Chen, Luyu, et al.
Pubblicazione: (2025)
di: Chen, Luyu, et al.
Pubblicazione: (2025)
Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction
di: Sheng, Huanxin, et al.
Pubblicazione: (2025)
di: Sheng, Huanxin, et al.
Pubblicazione: (2025)
Improving LLM-as-a-Judge Inference with the Judgment Distribution
di: Wang, Victor, et al.
Pubblicazione: (2025)
di: Wang, Victor, et al.
Pubblicazione: (2025)
Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces
di: Zhang, Chenchen
Pubblicazione: (2026)
di: Zhang, Chenchen
Pubblicazione: (2026)
Reinforcement Retrieval Leveraging Fine-grained Feedback for Fact Checking News Claims with Black-Box LLM
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
di: Zhou, Xin, et al.
Pubblicazione: (2025)
di: Zhou, Xin, et al.
Pubblicazione: (2025)
Can LLM be a Personalized Judge?
di: Dong, Yijiang River, et al.
Pubblicazione: (2024)
di: Dong, Yijiang River, et al.
Pubblicazione: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
di: Shi, Lin, et al.
Pubblicazione: (2024)
di: Shi, Lin, et al.
Pubblicazione: (2024)
Offline Reinforcement Learning for LLM Multi-Step Reasoning
di: Wang, Huaijie, et al.
Pubblicazione: (2024)
di: Wang, Huaijie, et al.
Pubblicazione: (2024)
Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
di: Lai, Peng, et al.
Pubblicazione: (2025)
di: Lai, Peng, et al.
Pubblicazione: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
Speech as a Multimodal Digital Phenotype for Multi-Task LLM-based Mental Health Prediction
di: Ali, Mai, et al.
Pubblicazione: (2025)
di: Ali, Mai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Bi-Level Prompt Optimization for Multimodal LLM-as-a-Judge
di: Pan, Bo, et al.
Pubblicazione: (2026) -
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
di: Jiang, Hongchao, et al.
Pubblicazione: (2025) -
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
di: Chen, Dongping, et al.
Pubblicazione: (2024) -
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
di: Wan, Zhongwei, et al.
Pubblicazione: (2025) -
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
di: Yang, Bo, et al.
Pubblicazione: (2026)