Bi-Level Prompt Optimization for Multimodal LLM-as-a-Judge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Bo, Kan, Xuan, Zhang, Kaitai, Yan, Yan, Tan, Shunwen, He, Zihao, Ding, Zixin, Wu, Junjie, Zhao, Liang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge
von: Wu, Junjie, et al.
Veröffentlicht: (2026)
von: Wu, Junjie, et al.
Veröffentlicht: (2026)
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
von: Pan, Tianjun, et al.
Veröffentlicht: (2026)
von: Pan, Tianjun, et al.
Veröffentlicht: (2026)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
von: Shi, Jiawen, et al.
Veröffentlicht: (2024)
von: Shi, Jiawen, et al.
Veröffentlicht: (2024)
Auto-Prompt Ensemble for LLM Judge
von: Li, Jiajie, et al.
Veröffentlicht: (2025)
von: Li, Jiajie, et al.
Veröffentlicht: (2025)
Can Past Experience Accelerate LLM Reasoning?
von: Pan, Bo, et al.
Veröffentlicht: (2025)
von: Pan, Bo, et al.
Veröffentlicht: (2025)
Bi-Level Optimization for Single Domain Generalization
von: Heidari, Marzi, et al.
Veröffentlicht: (2026)
von: Heidari, Marzi, et al.
Veröffentlicht: (2026)
Instructional Prompt Optimization for Few-Shot LLM-Based Recommendations on Cold-Start Users
von: Yang, Haowei, et al.
Veröffentlicht: (2025)
von: Yang, Haowei, et al.
Veröffentlicht: (2025)
Who's Your Judge? On the Detectability of LLM-Generated Judgments
von: Li, Dawei, et al.
Veröffentlicht: (2025)
von: Li, Dawei, et al.
Veröffentlicht: (2025)
Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
von: Koo, Hamin, et al.
Veröffentlicht: (2025)
von: Koo, Hamin, et al.
Veröffentlicht: (2025)
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
von: Han, Chao, et al.
Veröffentlicht: (2025)
von: Han, Chao, et al.
Veröffentlicht: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024)
von: Shi, Lin, et al.
Veröffentlicht: (2024)
JudgeFlow: Agentic Workflow Optimization via Block Judge
von: Ma, Zihan, et al.
Veröffentlicht: (2026)
von: Ma, Zihan, et al.
Veröffentlicht: (2026)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplifications and Resistance in Multi-Agent Based LLM-as-Judge
von: Ma, Chiyu, et al.
Veröffentlicht: (2025)
von: Ma, Chiyu, et al.
Veröffentlicht: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
LLM and GNN are Complementary: Distilling LLM for Multimodal Graph Learning
von: Xu, Junjie, et al.
Veröffentlicht: (2024)
von: Xu, Junjie, et al.
Veröffentlicht: (2024)
A More Advanced Group Polarization Measurement Approach Based on LLM-Based Agents and Graphs
von: Liu, Zixin, et al.
Veröffentlicht: (2024)
von: Liu, Zixin, et al.
Veröffentlicht: (2024)
Unifying Temporal and Structural Credit Assignment in LLM-Based Multi-Agent Prompt Optimization
von: Li, Wenwu, et al.
Veröffentlicht: (2026)
von: Li, Wenwu, et al.
Veröffentlicht: (2026)
Multi-Agent Debate for LLM Judges with Adaptive Stability Detection
von: Hu, Tianyu, et al.
Veröffentlicht: (2025)
von: Hu, Tianyu, et al.
Veröffentlicht: (2025)
Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution
von: Tian, Zailong, et al.
Veröffentlicht: (2025)
von: Tian, Zailong, et al.
Veröffentlicht: (2025)
Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection
von: Pan, Junjun, et al.
Veröffentlicht: (2025)
von: Pan, Junjun, et al.
Veröffentlicht: (2025)
Bi-Level Policy Optimization with Nyström Hypergradients
von: Prakash, Arjun, et al.
Veröffentlicht: (2025)
von: Prakash, Arjun, et al.
Veröffentlicht: (2025)
An Adaptive Differentially Private Federated Learning Framework with Bi-level Optimization
von: Wang, Jin, et al.
Veröffentlicht: (2026)
von: Wang, Jin, et al.
Veröffentlicht: (2026)
BiVRec: Bidirectional View-based Multimodal Sequential Recommendation
von: Hu, Jiaxi, et al.
Veröffentlicht: (2024)
von: Hu, Jiaxi, et al.
Veröffentlicht: (2024)
Diagnosing Live Within-Policy Instruction Conflicts in LLM Agents with Witnessed Resolution Profiles
von: Yan, Lu, et al.
Veröffentlicht: (2026)
von: Yan, Lu, et al.
Veröffentlicht: (2026)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
Criterion Validity of LLM-as-Judge for Business Outcomes in Conversational Commerce
von: Chen, Liang, et al.
Veröffentlicht: (2026)
von: Chen, Liang, et al.
Veröffentlicht: (2026)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
BadPromptFL: A Novel Backdoor Threat to Prompt-based Federated Learning in Multimodal Models
von: Zhang, Maozhen, et al.
Veröffentlicht: (2025)
von: Zhang, Maozhen, et al.
Veröffentlicht: (2025)
Character-Level Perturbations Disrupt LLM Watermarks
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2025)
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2025)
JudgeSQL: Reasoning over SQL Candidates with Weighted Consensus Tournament
von: Bai, Jiayuan, et al.
Veröffentlicht: (2025)
von: Bai, Jiayuan, et al.
Veröffentlicht: (2025)
LLM-Assisted Op-Amp Behavioral-Level Design via Agentic Human-Mimicking Reasoning
von: Chen, Zihao, et al.
Veröffentlicht: (2026)
von: Chen, Zihao, et al.
Veröffentlicht: (2026)
LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection
von: Hossain, Akram, et al.
Veröffentlicht: (2026)
von: Hossain, Akram, et al.
Veröffentlicht: (2026)
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
von: Dev, Sunishchal, et al.
Veröffentlicht: (2026)
von: Dev, Sunishchal, et al.
Veröffentlicht: (2026)
Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation
von: Mahmood, Alhasan, et al.
Veröffentlicht: (2026)
von: Mahmood, Alhasan, et al.
Veröffentlicht: (2026)
Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
von: Soumik, Sadman Kabir
Veröffentlicht: (2026)
von: Soumik, Sadman Kabir
Veröffentlicht: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Adaptive Prompt Embedding Optimization for LLM Jailbreaking
von: Li, Miles Q., et al.
Veröffentlicht: (2026)
von: Li, Miles Q., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge
von: Wu, Junjie, et al.
Veröffentlicht: (2026) -
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
von: Pan, Tianjun, et al.
Veröffentlicht: (2026) -
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
von: Shi, Jiawen, et al.
Veröffentlicht: (2024) -
Auto-Prompt Ensemble for LLM Judge
von: Li, Jiajie, et al.
Veröffentlicht: (2025) -
Can Past Experience Accelerate LLM Reasoning?
von: Pan, Bo, et al.
Veröffentlicht: (2025)