Automated Concept Discovery for LLM-as-a-Judge Preference Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wedgwood, James, Yadav, Chhavi, Smith, Virginia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
von: Wedgwood, James, et al.
Veröffentlicht: (2026)
von: Wedgwood, James, et al.
Veröffentlicht: (2026)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
von: Yang, Jinming, et al.
Veröffentlicht: (2026)
von: Yang, Jinming, et al.
Veröffentlicht: (2026)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024)
von: Shi, Lin, et al.
Veröffentlicht: (2024)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
von: Han, Steve, et al.
Veröffentlicht: (2025)
von: Han, Steve, et al.
Veröffentlicht: (2025)
Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models
von: Kumar, Shachi H, et al.
Veröffentlicht: (2024)
von: Kumar, Shachi H, et al.
Veröffentlicht: (2024)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
A Survey on LLM-as-a-Judge
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
Curriculum Learning for Safety Alignment
von: Kumar, Sandeep, et al.
Veröffentlicht: (2026)
von: Kumar, Sandeep, et al.
Veröffentlicht: (2026)
Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
NOVA: An Agentic Framework for Automated Histopathology Analysis and Discovery
von: Vaidya, Anurag J., et al.
Veröffentlicht: (2025)
von: Vaidya, Anurag J., et al.
Veröffentlicht: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
von: Lai, Peng, et al.
Veröffentlicht: (2026)
von: Lai, Peng, et al.
Veröffentlicht: (2026)
Meta-Judging with Large Language Models: Concepts, Methods, and Challenges
von: Silva, Hugo, et al.
Veröffentlicht: (2026)
von: Silva, Hugo, et al.
Veröffentlicht: (2026)
LLM-as-a-Judge for Time Series Explanations
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
AutoQual: An LLM Agent for Automated Discovery of Interpretable Features for Review Quality Assessment
von: Lan, Xiaochong, et al.
Veröffentlicht: (2025)
von: Lan, Xiaochong, et al.
Veröffentlicht: (2025)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
von: Ye, Jiayi, et al.
Veröffentlicht: (2024)
von: Ye, Jiayi, et al.
Veröffentlicht: (2024)
Are We on the Right Way to Assessing LLM-as-a-Judge?
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2024)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2024)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
Think-J: Learning to Think for Generative LLM-as-a-Judge
von: Huang, Hui, et al.
Veröffentlicht: (2025)
von: Huang, Hui, et al.
Veröffentlicht: (2025)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
von: Enguehard, Joseph, et al.
Veröffentlicht: (2025)
von: Enguehard, Joseph, et al.
Veröffentlicht: (2025)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
Automating Intervention Discovery from Scientific Literature: A Progressive Ontology Prompting and Dual-LLM Framework
von: Hu, Yuting, et al.
Veröffentlicht: (2024)
von: Hu, Yuting, et al.
Veröffentlicht: (2024)
VERT: Reliable LLM Judges for Radiology Report Evaluation
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
The Challenges of Evaluating LLM Applications: An Analysis of Automated, Human, and LLM-Based Approaches
von: Abeysinghe, Bhashithe, et al.
Veröffentlicht: (2024)
von: Abeysinghe, Bhashithe, et al.
Veröffentlicht: (2024)
Dissecting Human and LLM Preferences
von: Li, Junlong, et al.
Veröffentlicht: (2024)
von: Li, Junlong, et al.
Veröffentlicht: (2024)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
von: Zhou, Xiaolin, et al.
Veröffentlicht: (2026)
von: Zhou, Xiaolin, et al.
Veröffentlicht: (2026)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
von: Fujinuma, Yoshinari
Veröffentlicht: (2025)
von: Fujinuma, Yoshinari
Veröffentlicht: (2025)
Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
von: Chen, Luyu, et al.
Veröffentlicht: (2025)
von: Chen, Luyu, et al.
Veröffentlicht: (2025)
Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2025)
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2025)
Approximating Human Preferences Using a Multi-Judge Learned System
von: Sprejer, Eitán, et al.
Veröffentlicht: (2025)
von: Sprejer, Eitán, et al.
Veröffentlicht: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Criterion Validity of LLM-as-Judge for Business Outcomes in Conversational Commerce
von: Chen, Liang, et al.
Veröffentlicht: (2026)
von: Chen, Liang, et al.
Veröffentlicht: (2026)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
von: Yuan, Tongxin, et al.
Veröffentlicht: (2024)
von: Yuan, Tongxin, et al.
Veröffentlicht: (2024)
M-Prometheus: A Suite of Open Multilingual LLM Judges
von: Pombal, José, et al.
Veröffentlicht: (2025)
von: Pombal, José, et al.
Veröffentlicht: (2025)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
von: Wedgwood, James, et al.
Veröffentlicht: (2026) -
Quantifying and Mitigating Self-Preference Bias of LLM Judges
von: Yang, Jinming, et al.
Veröffentlicht: (2026) -
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026) -
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024) -
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
von: Han, Steve, et al.
Veröffentlicht: (2025)