JudgeFlow: Agentic Workflow Optimization via Block Judge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Zihan, Zhao, Zhikai, Hua, Chuanbo, Berto, Federico, Park, Jinkyoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TrajEvo: Designing Trajectory Prediction Heuristics via LLM-driven Evolution
von: Zhao, Zhikai, et al.
Veröffentlicht: (2025)
von: Zhao, Zhikai, et al.
Veröffentlicht: (2025)
EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models
von: Zhao, Zhikai, et al.
Veröffentlicht: (2026)
von: Zhao, Zhikai, et al.
Veröffentlicht: (2026)
TrajEvo: Trajectory Prediction Heuristics Design via LLM-driven Evolution
von: Zhao, Zhikai, et al.
Veröffentlicht: (2025)
von: Zhao, Zhikai, et al.
Veröffentlicht: (2025)
RRNCO: Towards Real-World Routing with Neural Combinatorial Optimization
von: Son, Jiwoo, et al.
Veröffentlicht: (2025)
von: Son, Jiwoo, et al.
Veröffentlicht: (2025)
HiMAP: Learning Heuristics-Informed Policies for Large-Scale Multi-Agent Pathfinding
von: Tang, Huijie, et al.
Veröffentlicht: (2024)
von: Tang, Huijie, et al.
Veröffentlicht: (2024)
CAMP: Collaborative Attention Model with Profiles for Vehicle Routing Problems
von: Hua, Chuanbo, et al.
Veröffentlicht: (2025)
von: Hua, Chuanbo, et al.
Veröffentlicht: (2025)
PARCO: Parallel AutoRegressive Models for Multi-Agent Combinatorial Optimization
von: Berto, Federico, et al.
Veröffentlicht: (2024)
von: Berto, Federico, et al.
Veröffentlicht: (2024)
Ensembling Prioritized Hybrid Policies for Multi-agent Pathfinding
von: Tang, Huijie, et al.
Veröffentlicht: (2024)
von: Tang, Huijie, et al.
Veröffentlicht: (2024)
ReEvo: Large Language Models as Hyper-Heuristics with Reflective Evolution
von: Ye, Haoran, et al.
Veröffentlicht: (2024)
von: Ye, Haoran, et al.
Veröffentlicht: (2024)
RouteFinder: Towards Foundation Models for Vehicle Routing Problems
von: Berto, Federico, et al.
Veröffentlicht: (2024)
von: Berto, Federico, et al.
Veröffentlicht: (2024)
Rethinking Positional Encoding for Neural Vehicle Routing
von: Hua, Chuanbo, et al.
Veröffentlicht: (2026)
von: Hua, Chuanbo, et al.
Veröffentlicht: (2026)
USPR: Learning a Unified Solver for Profiled Routing
von: Hua, Chuanbo, et al.
Veröffentlicht: (2025)
von: Hua, Chuanbo, et al.
Veröffentlicht: (2025)
Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows
von: Ming, Yifei, et al.
Veröffentlicht: (2025)
von: Ming, Yifei, et al.
Veröffentlicht: (2025)
VRPAgent: LLM-Driven Discovery of Heuristic Operators for Vehicle Routing Problems
von: Hottung, André, et al.
Veröffentlicht: (2025)
von: Hottung, André, et al.
Veröffentlicht: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024)
von: Shi, Lin, et al.
Veröffentlicht: (2024)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
FormalJudge: A Neuro-Symbolic Paradigm for Agentic Oversight
von: Zhou, Jiayi, et al.
Veröffentlicht: (2026)
von: Zhou, Jiayi, et al.
Veröffentlicht: (2026)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
von: Choi, Junhyuk, et al.
Veröffentlicht: (2026)
von: Choi, Junhyuk, et al.
Veröffentlicht: (2026)
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
von: Soumik, Sadman Kabir
Veröffentlicht: (2026)
von: Soumik, Sadman Kabir
Veröffentlicht: (2026)
Understanding and Optimizing Agentic Workflows via Shapley value
von: Yang, Yingxuan, et al.
Veröffentlicht: (2025)
von: Yang, Yingxuan, et al.
Veröffentlicht: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection
von: Hossain, Akram, et al.
Veröffentlicht: (2026)
von: Hossain, Akram, et al.
Veröffentlicht: (2026)
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
von: Dev, Sunishchal, et al.
Veröffentlicht: (2026)
von: Dev, Sunishchal, et al.
Veröffentlicht: (2026)
Bi-Level Prompt Optimization for Multimodal LLM-as-a-Judge
von: Pan, Bo, et al.
Veröffentlicht: (2026)
von: Pan, Bo, et al.
Veröffentlicht: (2026)
Multi-Agent Dynamic Relational Reasoning for Social Robot Navigation
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
BuildEvo: Designing Building Energy Consumption Forecasting Heuristics via LLM-driven Evolution
von: Lin, Subin, et al.
Veröffentlicht: (2025)
von: Lin, Subin, et al.
Veröffentlicht: (2025)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows
von: Liu, Wei, et al.
Veröffentlicht: (2026)
von: Liu, Wei, et al.
Veröffentlicht: (2026)
JudgeLRM: Large Reasoning Models as a Judge
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation
von: Bhonsle, Roshita, et al.
Veröffentlicht: (2025)
von: Bhonsle, Roshita, et al.
Veröffentlicht: (2025)
How Trustworthy Are LLM-as-Judge Ratings for Interpretive Responses? Implications for Qualitative Research Workflows
von: Han, Songhee, et al.
Veröffentlicht: (2026)
von: Han, Songhee, et al.
Veröffentlicht: (2026)
DenoiseFlow: Uncertainty-Aware Denoising for Reliable LLM Agentic Workflows
von: Yan, Yandong, et al.
Veröffentlicht: (2026)
von: Yan, Yandong, et al.
Veröffentlicht: (2026)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
von: Lee, Sua, et al.
Veröffentlicht: (2026)
von: Lee, Sua, et al.
Veröffentlicht: (2026)
AdaptFlow: Adaptive Workflow Optimization via Meta-Learning
von: Zhu, Runchuan, et al.
Veröffentlicht: (2025)
von: Zhu, Runchuan, et al.
Veröffentlicht: (2025)
When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
von: Yu, Fangyi
Veröffentlicht: (2025)
von: Yu, Fangyi
Veröffentlicht: (2025)
Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplifications and Resistance in Multi-Agent Based LLM-as-Judge
von: Ma, Chiyu, et al.
Veröffentlicht: (2025)
von: Ma, Chiyu, et al.
Veröffentlicht: (2025)
Who's Your Judge? On the Detectability of LLM-Generated Judgments
von: Li, Dawei, et al.
Veröffentlicht: (2025)
von: Li, Dawei, et al.
Veröffentlicht: (2025)
PentestJudge: Judging Agent Behavior Against Operational Requirements
von: Caldwell, Shane, et al.
Veröffentlicht: (2025)
von: Caldwell, Shane, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TrajEvo: Designing Trajectory Prediction Heuristics via LLM-driven Evolution
von: Zhao, Zhikai, et al.
Veröffentlicht: (2025) -
EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models
von: Zhao, Zhikai, et al.
Veröffentlicht: (2026) -
TrajEvo: Trajectory Prediction Heuristics Design via LLM-driven Evolution
von: Zhao, Zhikai, et al.
Veröffentlicht: (2025) -
RRNCO: Towards Real-World Routing with Neural Combinatorial Optimization
von: Son, Jiwoo, et al.
Veröffentlicht: (2025) -
HiMAP: Learning Heuristics-Informed Policies for Large-Scale Multi-Agent Pathfinding
von: Tang, Huijie, et al.
Veröffentlicht: (2024)