LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
Fuente:
arXiv
Guardado en:
| Autores principales: | Vishnubhotla, Krishnapriya, Vajjala, Soumya, Vij, Akriti, Nejadgholi, Isar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
por: Nejadgholi, Isar, et al.
Publicado: (2026)
por: Nejadgholi, Isar, et al.
Publicado: (2026)
The GPT-WritingPrompts Dataset: A Comparative Analysis of Character Portrayal in Short Stories
por: Huang, Xi Yu, et al.
Publicado: (2024)
por: Huang, Xi Yu, et al.
Publicado: (2024)
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
por: Fraser, Kathleen C., et al.
Publicado: (2025)
por: Fraser, Kathleen C., et al.
Publicado: (2025)
The Emotion Dynamics of Literary Novels
por: Vishnubhotla, Krishnapriya, et al.
Publicado: (2024)
por: Vishnubhotla, Krishnapriya, et al.
Publicado: (2024)
WMT24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles
por: Dawkins, Hillary, et al.
Publicado: (2024)
por: Dawkins, Hillary, et al.
Publicado: (2024)
Gender-Neutral Machine Translation Strategies in Practice
por: Dawkins, Hillary, et al.
Publicado: (2025)
por: Dawkins, Hillary, et al.
Publicado: (2025)
The Problem with Safety Classification is not just the Models
por: Vajjala, Sowmya
Publicado: (2025)
por: Vajjala, Sowmya
Publicado: (2025)
Affect, Body, Cognition, Demographics, and Emotion: The ABCDE of Text Features for Computational Affective Science
por: Wahle, Jan Philip, et al.
Publicado: (2025)
por: Wahle, Jan Philip, et al.
Publicado: (2025)
Socially Aware Synthetic Data Generation for Suicidal Ideation Detection Using Large Language Models
por: Ghanadian, Hamideh, et al.
Publicado: (2024)
por: Ghanadian, Hamideh, et al.
Publicado: (2024)
The crime of being poor
por: Curto, Georgina, et al.
Publicado: (2023)
por: Curto, Georgina, et al.
Publicado: (2023)
Projective Methods for Mitigating Gender Bias in Pre-trained Language Models
por: Dawkins, Hillary, et al.
Publicado: (2024)
por: Dawkins, Hillary, et al.
Publicado: (2024)
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
por: Nejadgholi, Isar, et al.
Publicado: (2024)
por: Nejadgholi, Isar, et al.
Publicado: (2024)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
por: Nejadgholi, Isar, et al.
Publicado: (2025)
por: Nejadgholi, Isar, et al.
Publicado: (2025)
Emotion Granularity from Text: An Aggregate-Level Indicator of Mental Health
por: Vishnubhotla, Krishnapriya, et al.
Publicado: (2024)
por: Vishnubhotla, Krishnapriya, et al.
Publicado: (2024)
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse
por: Guo, Rongchen, et al.
Publicado: (2024)
por: Guo, Rongchen, et al.
Publicado: (2024)
Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles
por: Guo, Rongchen, et al.
Publicado: (2025)
por: Guo, Rongchen, et al.
Publicado: (2025)
Text Classification in the LLM Era -- Where do we stand?
por: Vajjala, Sowmya, et al.
Publicado: (2025)
por: Vajjala, Sowmya, et al.
Publicado: (2025)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
por: Wang, Yidong, et al.
Publicado: (2025)
por: Wang, Yidong, et al.
Publicado: (2025)
Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
por: Haldar, Rajarshi, et al.
Publicado: (2025)
por: Haldar, Rajarshi, et al.
Publicado: (2025)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
por: Sun, Bian, et al.
Publicado: (2026)
por: Sun, Bian, et al.
Publicado: (2026)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
por: Yang, Langqi, et al.
Publicado: (2025)
por: Yang, Langqi, et al.
Publicado: (2025)
IndicGEC: Powerful Models, or a Measurement Mirage?
por: Vajjala, Sowmya
Publicado: (2025)
por: Vajjala, Sowmya
Publicado: (2025)
Tackling Social Bias against the Poor: A Dataset and Taxonomy on Aporophobia
por: Curto, Georgina, et al.
Publicado: (2025)
por: Curto, Georgina, et al.
Publicado: (2025)
Same Input, Different Scores: A Multi Model Study on the Inconsistency of LLM Judge
por: Lau, Fiona
Publicado: (2026)
por: Lau, Fiona
Publicado: (2026)
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
por: Subramonian, Arjun, et al.
Publicado: (2025)
por: Subramonian, Arjun, et al.
Publicado: (2025)
Dravidian language family through Universal Dependencies lens
por: Rama, Taraka, et al.
Publicado: (2024)
por: Rama, Taraka, et al.
Publicado: (2024)
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
Does Synthetic Data Help Named Entity Recognition for Low-Resource Languages?
por: Kamath, Gaurav, et al.
Publicado: (2025)
por: Kamath, Gaurav, et al.
Publicado: (2025)
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
Should LLM Safety Be More Than Refusing Harmful Instructions?
por: Maskey, Utsav, et al.
Publicado: (2025)
por: Maskey, Utsav, et al.
Publicado: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
por: Clegg, Kester, et al.
Publicado: (2025)
por: Clegg, Kester, et al.
Publicado: (2025)
Test Set Quality in Multilingual LLM Evaluation
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
Annotation Errors and NER: A Study with OntoNotes 5.0
por: Bernier-Colborne, Gabriel, et al.
Publicado: (2024)
por: Bernier-Colborne, Gabriel, et al.
Publicado: (2024)
Social and Ethical Risks Posed by General-Purpose LLMs for Settling Newcomers in Canada
por: Nejadgholi, Isar, et al.
Publicado: (2024)
por: Nejadgholi, Isar, et al.
Publicado: (2024)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
por: Zhang, Chi, et al.
Publicado: (2025)
por: Zhang, Chi, et al.
Publicado: (2025)
Modeling Contextual Passage Utility for Multihop Question Answering
por: Jain, Akriti, et al.
Publicado: (2025)
por: Jain, Akriti, et al.
Publicado: (2025)
Knowing What's Missing: Assessing Information Sufficiency in Question Answering
por: Jain, Akriti, et al.
Publicado: (2025)
por: Jain, Akriti, et al.
Publicado: (2025)
Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing
por: Ahn, Jihyun Janice, et al.
Publicado: (2025)
por: Ahn, Jihyun Janice, et al.
Publicado: (2025)
Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties
por: Faisal, Fahim, et al.
Publicado: (2024)
por: Faisal, Fahim, et al.
Publicado: (2024)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
por: Yuan, Tongxin, et al.
Publicado: (2024)
por: Yuan, Tongxin, et al.
Publicado: (2024)
Ejemplares similares
-
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
por: Nejadgholi, Isar, et al.
Publicado: (2026) -
The GPT-WritingPrompts Dataset: A Comparative Analysis of Character Portrayal in Short Stories
por: Huang, Xi Yu, et al.
Publicado: (2024) -
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
por: Fraser, Kathleen C., et al.
Publicado: (2025) -
The Emotion Dynamics of Literary Novels
por: Vishnubhotla, Krishnapriya, et al.
Publicado: (2024) -
WMT24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles
por: Dawkins, Hillary, et al.
Publicado: (2024)