HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Langqi, Zheng, Tianhang, Chen, Yixuan, Xiu, Kedong, Zhou, Hao, Ni, Wangze, Chen, Lei, Qin, Zhan, Ren, Kui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
por: Huang, Xinzhe, et al.
Publicado: (2025)
por: Huang, Xinzhe, et al.
Publicado: (2025)
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
por: Yan, Jianxin, et al.
Publicado: (2025)
por: Yan, Jianxin, et al.
Publicado: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
por: Vishnubhotla, Krishnapriya, et al.
Publicado: (2026)
por: Vishnubhotla, Krishnapriya, et al.
Publicado: (2026)
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
por: Cyberey, Hannah, et al.
Publicado: (2024)
por: Cyberey, Hannah, et al.
Publicado: (2024)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
por: Zhou, Xin, et al.
Publicado: (2025)
por: Zhou, Xin, et al.
Publicado: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
por: Clegg, Kester, et al.
Publicado: (2025)
por: Clegg, Kester, et al.
Publicado: (2025)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
por: Liu, Kangwei, et al.
Publicado: (2025)
por: Liu, Kangwei, et al.
Publicado: (2025)
Self-HarmLLM: Can Large Language Model Harm Itself?
por: Kim, Heehwan, et al.
Publicado: (2025)
por: Kim, Heehwan, et al.
Publicado: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
por: Li, Jing-Jing, et al.
Publicado: (2026)
por: Li, Jing-Jing, et al.
Publicado: (2026)
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
por: Sharshar, Ahmed, et al.
Publicado: (2026)
por: Sharshar, Ahmed, et al.
Publicado: (2026)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
por: Pandey, Punya Syon, et al.
Publicado: (2025)
por: Pandey, Punya Syon, et al.
Publicado: (2025)
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
por: Zhou, Yuhang, et al.
Publicado: (2025)
por: Zhou, Yuhang, et al.
Publicado: (2025)
Untargeted Jailbreak Attack
por: Huang, Xinzhe, et al.
Publicado: (2025)
por: Huang, Xinzhe, et al.
Publicado: (2025)
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
por: Zeng, Churui, et al.
Publicado: (2026)
por: Zeng, Churui, et al.
Publicado: (2026)
RepEval: Effective Text Evaluation with LLM Representation
por: Sheng, Shuqian, et al.
Publicado: (2024)
por: Sheng, Shuqian, et al.
Publicado: (2024)
Can Editing LLMs Inject Harm?
por: Chen, Canyu, et al.
Publicado: (2024)
por: Chen, Canyu, et al.
Publicado: (2024)
Investigating and Alleviating Harm Amplification in LLM Interactions
por: Guo, Ruohao, et al.
Publicado: (2026)
por: Guo, Ruohao, et al.
Publicado: (2026)
Deep Research Brings Deeper Harm
por: Chen, Shuo, et al.
Publicado: (2025)
por: Chen, Shuo, et al.
Publicado: (2025)
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
por: Eigler, Lukáš, et al.
Publicado: (2026)
por: Eigler, Lukáš, et al.
Publicado: (2026)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
por: Zhou, Yuhang, et al.
Publicado: (2025)
por: Zhou, Yuhang, et al.
Publicado: (2025)
LLM-based Semantic Augmentation for Harmful Content Detection
por: Meguellati, Elyas, et al.
Publicado: (2025)
por: Meguellati, Elyas, et al.
Publicado: (2025)
Dynamic Target Attack
por: Xiu, Kedong, et al.
Publicado: (2025)
por: Xiu, Kedong, et al.
Publicado: (2025)
Should LLM Safety Be More Than Refusing Harmful Instructions?
por: Maskey, Utsav, et al.
Publicado: (2025)
por: Maskey, Utsav, et al.
Publicado: (2025)
HarmPot: An Annotation Framework for Evaluating Offline Harm Potential of Social Media Text
por: Kumar, Ritesh, et al.
Publicado: (2024)
por: Kumar, Ritesh, et al.
Publicado: (2024)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
por: Zhou, Lingfeng, et al.
Publicado: (2025)
por: Zhou, Lingfeng, et al.
Publicado: (2025)
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
por: Chen, Weiyi, et al.
Publicado: (2026)
por: Chen, Weiyi, et al.
Publicado: (2026)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
por: Rao, Delip, et al.
Publicado: (2026)
por: Rao, Delip, et al.
Publicado: (2026)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
por: Son, Guijin, et al.
Publicado: (2024)
por: Son, Guijin, et al.
Publicado: (2024)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
por: Wei, Hui, et al.
Publicado: (2024)
por: Wei, Hui, et al.
Publicado: (2024)
When Harmless Words Harm: A New Threat to LLM Safety via Conceptual Triggers
por: Zhang, Zhaoxin, et al.
Publicado: (2025)
por: Zhang, Zhaoxin, et al.
Publicado: (2025)
LecEval: An Automated Metric for Multimodal Knowledge Acquisition in Multimedia Learning
por: Yin, Joy Lim Jia, et al.
Publicado: (2025)
por: Yin, Joy Lim Jia, et al.
Publicado: (2025)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
por: Mekky, Ali, et al.
Publicado: (2025)
por: Mekky, Ali, et al.
Publicado: (2025)
Representational Harms in LLM-Generated Narratives Against Global Majority Nationalities
por: Nguyen, Ilana, et al.
Publicado: (2026)
por: Nguyen, Ilana, et al.
Publicado: (2026)
HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate
por: Zhu, Shenzhe
Publicado: (2025)
por: Zhu, Shenzhe
Publicado: (2025)
HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment
por: Belkhiter, Yannis, et al.
Publicado: (2024)
por: Belkhiter, Yannis, et al.
Publicado: (2024)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
por: Alam, Firoj, et al.
Publicado: (2026)
por: Alam, Firoj, et al.
Publicado: (2026)
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?
por: Chen, Jiamin, et al.
Publicado: (2026)
por: Chen, Jiamin, et al.
Publicado: (2026)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
por: Barnes, Jeremy, et al.
Publicado: (2025)
por: Barnes, Jeremy, et al.
Publicado: (2025)
RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
por: Zhang, Qiyuan, et al.
Publicado: (2024)
por: Zhang, Qiyuan, et al.
Publicado: (2024)
Ejemplares similares
-
DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
por: Huang, Xinzhe, et al.
Publicado: (2025) -
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
por: Yan, Jianxin, et al.
Publicado: (2025) -
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
por: Andriushchenko, Maksym, et al.
Publicado: (2024) -
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
por: Vishnubhotla, Krishnapriya, et al.
Publicado: (2026) -
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
por: Cyberey, Hannah, et al.
Publicado: (2024)