Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
Fuente:
arXiv
Saved in:
| Main Authors: | Gajcin, Jasmina, Miehling, Erik, Nair, Rahul, Daly, Elizabeth, Marinescu, Radu, Tirupathi, Seshu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Who Sees the Risk? Stakeholder Conflicts and Explanatory Policies in LLM-based Risk Assessment
by: Yadav, Srishti, et al.
Published: (2025)
by: Yadav, Srishti, et al.
Published: (2025)
FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models
by: Carnerero-Cano, Javier, et al.
Published: (2026)
by: Carnerero-Cano, Javier, et al.
Published: (2026)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
by: Do, Hyo Jin, et al.
Published: (2025)
by: Do, Hyo Jin, et al.
Published: (2025)
Localizing Persona Representations in LLMs
by: Cintas, Celia, et al.
Published: (2025)
by: Cintas, Celia, et al.
Published: (2025)
Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring
by: Belkhiter, Yannis, et al.
Published: (2025)
by: Belkhiter, Yannis, et al.
Published: (2025)
Redefining Counterfactual Explanations for Reinforcement Learning: Overview, Challenges and Opportunities
by: Gajcin, Jasmina, et al.
Published: (2022)
by: Gajcin, Jasmina, et al.
Published: (2022)
CELL your Model: Contrastive Explanations for Large Language Models
by: Luss, Ronny, et al.
Published: (2024)
by: Luss, Ronny, et al.
Published: (2024)
GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models
by: Tirupathi, Seshu, et al.
Published: (2025)
by: Tirupathi, Seshu, et al.
Published: (2025)
Semifactual Explanations for Reinforcement Learning
by: Gajcin, Jasmina, et al.
Published: (2024)
by: Gajcin, Jasmina, et al.
Published: (2024)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models
by: Belkhiter, Yannis, et al.
Published: (2026)
by: Belkhiter, Yannis, et al.
Published: (2026)
ACTER: Diverse and Actionable Counterfactual Sequences for Explaining and Diagnosing RL Policies
by: Gajcin, Jasmina, et al.
Published: (2024)
by: Gajcin, Jasmina, et al.
Published: (2024)
LLM-as-a-Judge for Time Series Explanations
by: Sivalingam, Preetham, et al.
Published: (2026)
by: Sivalingam, Preetham, et al.
Published: (2026)
FactReasoner: A Probabilistic Approach to Long-Form Factuality Assessment for Large Language Models
by: Marinescu, Radu, et al.
Published: (2025)
by: Marinescu, Radu, et al.
Published: (2025)
Human and Automatic Interpretation of Romanian Noun Compounds
by: Marinescu, Ioana, et al.
Published: (2024)
by: Marinescu, Ioana, et al.
Published: (2024)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
by: Hou, Yufang, et al.
Published: (2024)
by: Hou, Yufang, et al.
Published: (2024)
Ranking Large Language Models without Ground Truth
by: Dhurandhar, Amit, et al.
Published: (2024)
by: Dhurandhar, Amit, et al.
Published: (2024)
Evaluating the Prompt Steerability of Large Language Models
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
by: Zhang, Taolin, et al.
Published: (2025)
by: Zhang, Taolin, et al.
Published: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
by: Liu, Yixin, et al.
Published: (2026)
by: Liu, Yixin, et al.
Published: (2026)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
by: Shi, Lin, et al.
Published: (2024)
by: Shi, Lin, et al.
Published: (2024)
How Trustworthy Are LLM-as-Judge Ratings for Interpretive Responses? Implications for Qualitative Research Workflows
by: Han, Songhee, et al.
Published: (2026)
by: Han, Songhee, et al.
Published: (2026)
Multimodal Analysis of State-Funded News Coverage of the Israel-Hamas War on YouTube Shorts
by: Miehling, Daniel, et al.
Published: (2026)
by: Miehling, Daniel, et al.
Published: (2026)
AI Steerability 360: A Toolkit for Steering Large Language Models
by: Miehling, Erik, et al.
Published: (2026)
by: Miehling, Erik, et al.
Published: (2026)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
by: Wang, Yidong, et al.
Published: (2025)
by: Wang, Yidong, et al.
Published: (2025)
Vichara: Appellate Judgment Prediction and Explanation for the Indian Judicial System
by: Nair, Pavithra PM, et al.
Published: (2026)
by: Nair, Pavithra PM, et al.
Published: (2026)
A Survey on LLM-as-a-Judge
by: Gu, Jiawei, et al.
Published: (2024)
by: Gu, Jiawei, et al.
Published: (2024)
SIMBA UQ: Similarity-Based Aggregation for Uncertainty Quantification in Large Language Models
by: Bhattacharjya, Debarun, et al.
Published: (2025)
by: Bhattacharjya, Debarun, et al.
Published: (2025)
EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
Automatic Replication of LLM Mistakes in Medical Conversations
by: Proniakin, Oleksii, et al.
Published: (2025)
by: Proniakin, Oleksii, et al.
Published: (2025)
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
by: Zhang, Hongbin, et al.
Published: (2026)
by: Zhang, Hongbin, et al.
Published: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
by: Clegg, Kester, et al.
Published: (2025)
by: Clegg, Kester, et al.
Published: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
by: Tong, Terry, et al.
Published: (2025)
by: Tong, Terry, et al.
Published: (2025)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
by: Ye, Jiayi, et al.
Published: (2024)
by: Ye, Jiayi, et al.
Published: (2024)
Are We on the Right Way to Assessing LLM-as-a-Judge?
by: Feng, Yuanning, et al.
Published: (2025)
by: Feng, Yuanning, et al.
Published: (2025)
Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
by: McCarthy, James, et al.
Published: (2025)
by: McCarthy, James, et al.
Published: (2025)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
by: Han, Steve, et al.
Published: (2025)
by: Han, Steve, et al.
Published: (2025)
JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
Support-Contra Asymmetry in LLM Explanations
by: Patil, Avinash
Published: (2025)
by: Patil, Avinash
Published: (2025)
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier
by: Jiang, Yuhua, et al.
Published: (2025)
by: Jiang, Yuhua, et al.
Published: (2025)
Similar Items
-
Who Sees the Risk? Stakeholder Conflicts and Explanatory Policies in LLM-based Risk Assessment
by: Yadav, Srishti, et al.
Published: (2025) -
FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models
by: Carnerero-Cano, Javier, et al.
Published: (2026) -
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
by: Do, Hyo Jin, et al.
Published: (2025) -
Localizing Persona Representations in LLMs
by: Cintas, Celia, et al.
Published: (2025) -
Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring
by: Belkhiter, Yannis, et al.
Published: (2025)