LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
Fuente:
arXiv
Guardado en:
| Autores principales: | Park, Junyeong, Jeong, Seogyeong, Song, Seyoung, Lee, Yohan, Oh, Alice |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
por: Song, Seyoung, et al.
Publicado: (2025)
por: Song, Seyoung, et al.
Publicado: (2025)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
por: Jin, Jiho, et al.
Publicado: (2026)
por: Jin, Jiho, et al.
Publicado: (2026)
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
por: Lee, Nayeon, et al.
Publicado: (2023)
por: Lee, Nayeon, et al.
Publicado: (2023)
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
por: Yu, Haeun, et al.
Publicado: (2025)
por: Yu, Haeun, et al.
Publicado: (2025)
Human Psychometric Questionnaires Mischaracterize LLM Behavior
por: Song, Woojung, et al.
Publicado: (2025)
por: Song, Woojung, et al.
Publicado: (2025)
LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control
por: Jeong, Seogyeong, et al.
Publicado: (2026)
por: Jeong, Seogyeong, et al.
Publicado: (2026)
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
por: Giorgi, Tommaso, et al.
Publicado: (2024)
por: Giorgi, Tommaso, et al.
Publicado: (2024)
XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
por: Choi, Dasol, et al.
Publicado: (2026)
por: Choi, Dasol, et al.
Publicado: (2026)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
por: Piot, Paloma, et al.
Publicado: (2025)
por: Piot, Paloma, et al.
Publicado: (2025)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
por: Hu, Yujia, et al.
Publicado: (2026)
por: Hu, Yujia, et al.
Publicado: (2026)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
por: Kumar, Shanu, et al.
Publicado: (2024)
por: Kumar, Shanu, et al.
Publicado: (2024)
Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
por: Wang, Yifan, et al.
Publicado: (2025)
por: Wang, Yifan, et al.
Publicado: (2025)
Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection
por: Kim, Yumin, et al.
Publicado: (2025)
por: Kim, Yumin, et al.
Publicado: (2025)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
por: Kim, Eunsu, et al.
Publicado: (2024)
por: Kim, Eunsu, et al.
Publicado: (2024)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
por: Proskurina, Irina, et al.
Publicado: (2025)
por: Proskurina, Irina, et al.
Publicado: (2025)
LLM-Human Pipeline for Cultural Context Grounding of Conversations
por: Pujari, Rajkumar, et al.
Publicado: (2024)
por: Pujari, Rajkumar, et al.
Publicado: (2024)
Learning to Summarize from LLM-generated Feedback
por: Song, Hwanjun, et al.
Publicado: (2024)
por: Song, Hwanjun, et al.
Publicado: (2024)
Causality Guided Representation Learning for Cross-Style Hate Speech Detection
por: Zhao, Chengshuai, et al.
Publicado: (2025)
por: Zhao, Chengshuai, et al.
Publicado: (2025)
NoisyHate: Mining Online Human-Written Perturbations for Realistic Robustness Benchmarking of Content Moderation Models
por: Ye, Yiran, et al.
Publicado: (2023)
por: Ye, Yiran, et al.
Publicado: (2023)
Empirical Evaluation of Public HateSpeech Datasets
por: Jaf, Sadar, et al.
Publicado: (2024)
por: Jaf, Sadar, et al.
Publicado: (2024)
Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech
por: Ludwig, Florian, et al.
Publicado: (2025)
por: Ludwig, Florian, et al.
Publicado: (2025)
xList-Hate: A Checklist-Based Framework for Interpretable and Generalizable Hate Speech Detection
por: Girón, Adrián, et al.
Publicado: (2026)
por: Girón, Adrián, et al.
Publicado: (2026)
An Investigation Into Explainable Audio Hate Speech Detection
por: An, Jinmyeong, et al.
Publicado: (2024)
por: An, Jinmyeong, et al.
Publicado: (2024)
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
por: Seo, Gyuhyeon, et al.
Publicado: (2025)
por: Seo, Gyuhyeon, et al.
Publicado: (2025)
Collaboratively adding new knowledge to an LLM
por: Lee, Rhui Dih, et al.
Publicado: (2024)
por: Lee, Rhui Dih, et al.
Publicado: (2024)
Probing Association Biases in LLM Moderation Over-Sensitivity
por: Wang, Yuxin, et al.
Publicado: (2025)
por: Wang, Yuxin, et al.
Publicado: (2025)
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
por: Wu, Wanxing, et al.
Publicado: (2026)
por: Wu, Wanxing, et al.
Publicado: (2026)
Improving LLM Classification of Logical Errors by Integrating Error Relationship into Prompts
por: Lee, Yanggyu, et al.
Publicado: (2024)
por: Lee, Yanggyu, et al.
Publicado: (2024)
Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
por: Saha, Sougata, et al.
Publicado: (2024)
por: Saha, Sougata, et al.
Publicado: (2024)
SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia
por: Ng, Ri Chi, et al.
Publicado: (2026)
por: Ng, Ri Chi, et al.
Publicado: (2026)
Improving Hate Speech Classification with Cross-Taxonomy Dataset Integration
por: Fillies, Jan, et al.
Publicado: (2025)
por: Fillies, Jan, et al.
Publicado: (2025)
BengaliSent140: A Large-Scale Bengali Binary Sentiment Dataset for Hate and Non-Hate Speech Classification
por: Islam, Akif, et al.
Publicado: (2026)
por: Islam, Akif, et al.
Publicado: (2026)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
por: Kumarage, Tharindu, et al.
Publicado: (2024)
por: Kumarage, Tharindu, et al.
Publicado: (2024)
CollabStory: Multi-LLM Collaborative Story Generation and Authorship Analysis
por: Venkatraman, Saranya, et al.
Publicado: (2024)
por: Venkatraman, Saranya, et al.
Publicado: (2024)
Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
por: Selvaganapathy, Sanjeeevan, et al.
Publicado: (2025)
por: Selvaganapathy, Sanjeeevan, et al.
Publicado: (2025)
A Cross-Cultural Assessment of Human Ability to Detect LLM-Generated Fake News about South Africa
por: Schlippe, Tim, et al.
Publicado: (2025)
por: Schlippe, Tim, et al.
Publicado: (2025)
Cross-Task Experiential Learning on LLM-based Multi-Agent Collaboration
por: Li, Yilong, et al.
Publicado: (2025)
por: Li, Yilong, et al.
Publicado: (2025)
FLAME: Flexible LLM-Assisted Moderation Engine
por: Bakulin, Ivan, et al.
Publicado: (2025)
por: Bakulin, Ivan, et al.
Publicado: (2025)
Serendipity by Design: Evaluating the Impact of Cross-domain Mappings on Human and LLM Creativity
por: Liu, Qiawen Ella, et al.
Publicado: (2026)
por: Liu, Qiawen Ella, et al.
Publicado: (2026)
MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents
por: Lei, Yuzhen, et al.
Publicado: (2025)
por: Lei, Yuzhen, et al.
Publicado: (2025)
Ejemplares similares
-
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
por: Song, Seyoung, et al.
Publicado: (2025) -
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
por: Jin, Jiho, et al.
Publicado: (2026) -
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
por: Lee, Nayeon, et al.
Publicado: (2023) -
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
por: Yu, Haeun, et al.
Publicado: (2025) -
Human Psychometric Questionnaires Mischaracterize LLM Behavior
por: Song, Woojung, et al.
Publicado: (2025)