Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Leyao, He, Yanan, Chen, Peng, Yehudai, Asaf, Liu, Yixin, Ying, Rex, Shmueli-Scheuer, Michal, Cohan, Arman |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
by: Yehudai, Asaf, et al.
Published: (2026)
by: Yehudai, Asaf, et al.
Published: (2026)
Survey on Evaluation of LLM-based Agents
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
by: Keren, Tomer, et al.
Published: (2026)
by: Keren, Tomer, et al.
Published: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
by: Perlitz, Yotam, et al.
Published: (2024)
by: Perlitz, Yotam, et al.
Published: (2024)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
by: Habba, Eliya, et al.
Published: (2026)
by: Habba, Eliya, et al.
Published: (2026)
Mediocrity is the key for LLM as a Judge Anchor Selection
by: Don-Yehiya, Shachar, et al.
Published: (2026)
by: Don-Yehiya, Shachar, et al.
Published: (2026)
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
by: Kour, George, et al.
Published: (2025)
by: Kour, George, et al.
Published: (2025)
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024)
by: Gera, Ariel, et al.
Published: (2024)
Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction
by: Bui, Ngoc, et al.
Published: (2026)
by: Bui, Ngoc, et al.
Published: (2026)
General Agent Evaluation
by: Bandel, Elron, et al.
Published: (2026)
by: Bandel, Elron, et al.
Published: (2026)
ResearchGym: Evaluating Language Model Agents on Real-World AI Research
by: Garikaparthi, Aniketh, et al.
Published: (2026)
by: Garikaparthi, Aniketh, et al.
Published: (2026)
Understanding Reference Policies in Direct Preference Optimization
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes
by: Yehudai, Asaf, et al.
Published: (2024)
by: Yehudai, Asaf, et al.
Published: (2024)
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
by: Ashury-Tahan, Shir, et al.
Published: (2025)
by: Ashury-Tahan, Shir, et al.
Published: (2025)
Benchmark Leakage Trap: Can We Trust LLM-based Recommendation?
by: Zhang, Mingqiao, et al.
Published: (2026)
by: Zhang, Mingqiao, et al.
Published: (2026)
Can We Trust LLM Detectors?
by: Sandhan, Jivnesh, et al.
Published: (2026)
by: Sandhan, Jivnesh, et al.
Published: (2026)
LLM-REVal: Can We Trust LLM Reviewers Yet?
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
by: Liu, Yixin, et al.
Published: (2026)
by: Liu, Yixin, et al.
Published: (2026)
Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplifications and Resistance in Multi-Agent Based LLM-as-Judge
by: Ma, Chiyu, et al.
Published: (2025)
by: Ma, Chiyu, et al.
Published: (2025)
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
by: Schroeder, Kayla, et al.
Published: (2024)
by: Schroeder, Kayla, et al.
Published: (2024)
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
by: Yehudai, Asaf, et al.
Published: (2026)
by: Yehudai, Asaf, et al.
Published: (2026)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
by: Xu, Zhijian, et al.
Published: (2025)
by: Xu, Zhijian, et al.
Published: (2025)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
Robustness as an Emergent Property of Task Performance
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
Can We Trust AI to Govern AI? Benchmarking LLM Performance on Privacy and AI Governance Exams
by: Witherspoon, Zane, et al.
Published: (2025)
by: Witherspoon, Zane, et al.
Published: (2025)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
by: Hu, Tiansheng, et al.
Published: (2026)
by: Hu, Tiansheng, et al.
Published: (2026)
References Improve LLM Alignment in Non-Verifiable Domains
by: Shi, Kejian, et al.
Published: (2026)
by: Shi, Kejian, et al.
Published: (2026)
SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking
by: Li, Jindong, et al.
Published: (2026)
by: Li, Jindong, et al.
Published: (2026)
YaleNLP @ PerAnsSumm 2025: Multi-Perspective Integration via Mixture-of-Agents for Enhanced Healthcare QA Summarization
by: Jang, Dongsuk, et al.
Published: (2025)
by: Jang, Dongsuk, et al.
Published: (2025)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
Observable Propagation: Uncovering Feature Vectors in Transformers
by: Dunefsky, Jacob, et al.
Published: (2023)
by: Dunefsky, Jacob, et al.
Published: (2023)
One-shot Optimized Steering Vectors Mediate Safety-relevant Behaviors in LLMs
by: Dunefsky, Jacob, et al.
Published: (2025)
by: Dunefsky, Jacob, et al.
Published: (2025)
Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-based Decision-Making Systems
by: Jiao, Ruochen, et al.
Published: (2024)
by: Jiao, Ruochen, et al.
Published: (2024)
Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
by: Uzan, Omri, et al.
Published: (2025)
by: Uzan, Omri, et al.
Published: (2025)
Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
by: Li, Alan, et al.
Published: (2025)
by: Li, Alan, et al.
Published: (2025)
Calibrating Long-form Generations from Large Language Models
by: Huang, Yukun, et al.
Published: (2024)
by: Huang, Yukun, et al.
Published: (2024)
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Similar Items
-
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
by: Yehudai, Asaf, et al.
Published: (2026) -
Survey on Evaluation of LLM-based Agents
by: Yehudai, Asaf, et al.
Published: (2025) -
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
by: Yehudai, Asaf, et al.
Published: (2025) -
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
by: Keren, Tomer, et al.
Published: (2026) -
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)