Supporting Human Raters with the Detection of Harmful Content using Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Thomas, Kurt, Kelley, Patrick Gage, Tao, David, Meiklejohn, Sarah, Vallis, Owen, Tan, Shunwen, Bratanič, Blaž, Ferreira, Felipe Tiengo, Eranti, Vijay Kumar, Bursztein, Elie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RETVec: Resilient and Efficient Text Vectorizer
by: Bursztein, Elie, et al.
Published: (2023)
by: Bursztein, Elie, et al.
Published: (2023)
Understanding Help-Seeking and Help-Giving on Social Media for Image-Based Sexual Abuse
by: Wei, Miranda, et al.
Published: (2024)
by: Wei, Miranda, et al.
Published: (2024)
"It didn't feel right but I needed a job so desperately": Understanding People's Emotions & Help Needs During Financial Scams
by: Chanenson, Jake, et al.
Published: (2026)
by: Chanenson, Jake, et al.
Published: (2026)
Understanding Help Seeking for Digital Privacy, Safety, and Security
by: Thomas, Kurt, et al.
Published: (2026)
by: Thomas, Kurt, et al.
Published: (2026)
Magika: AI-Powered Content-Type Detection
by: Fratantonio, Yanick, et al.
Published: (2024)
by: Fratantonio, Yanick, et al.
Published: (2024)
How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape
by: Kelley, Patrick Gage, et al.
Published: (2025)
by: Kelley, Patrick Gage, et al.
Published: (2025)
A Qualitative Content Analysis Exploring Systems‐Based Practice Across Health Professional Competency Standards
by: Claire Palermo, et al.
Published: (2026)
by: Claire Palermo, et al.
Published: (2026)
GASTROESOPHAGEAL REFLUX DISEASE: FROM PATHOPHYSIOLOGICAL MECHANISMS TO PERSONALIZED THERAPEUTIC STRATEGIES
by: Shaik Mahammad Shaheed, et al.
Published: (2026)
by: Shaik Mahammad Shaheed, et al.
Published: (2026)
"We are not Future-ready": Understanding AI Privacy Risks and Existing Mitigation Strategies from the Perspective of AI Developers in Europe
by: Klymenko, Alexandra, et al.
Published: (2025)
by: Klymenko, Alexandra, et al.
Published: (2025)
Privacy Risks of General-Purpose AI Systems: A Foundation for Investigating Practitioner Perspectives
by: Meisenbacher, Stephen, et al.
Published: (2024)
by: Meisenbacher, Stephen, et al.
Published: (2024)
Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence
by: Resnick, Paul, et al.
Published: (2021)
by: Resnick, Paul, et al.
Published: (2021)
SYNAPSE-G: Bridging Large Language Models and Graph Learning for Rare Event Classification
by: Tavakkol, Sasan, et al.
Published: (2025)
by: Tavakkol, Sasan, et al.
Published: (2025)
Investigation of the Inter-Rater Reliability between Large Language Models and Human Raters in Qualitative Analysis
by: Borse, Nikhil Sanjay, et al.
Published: (2025)
by: Borse, Nikhil Sanjay, et al.
Published: (2025)
O Fenômeno População em Situação de Rua Enquanto Fruto do Capitalismo
by: Verônica Martins Tiengo
Published: (2018)
by: Verônica Martins Tiengo
Published: (2018)
Ação Integralista Brasileira (AIB) e Forças Armadas: notas de pesquisa através do jornal “Flamma Verde” (Florianópolis 1936-1938)
by: Gustavo Tiengo Pontes
Published: (2015)
by: Gustavo Tiengo Pontes
Published: (2015)
A PANDEMIA E SEUS IMPACTOS PARA A POPULAÇÃO EM SITUAÇÃO DE RUA
by: Verônica Martins Tiengo
Published: (2021)
by: Verônica Martins Tiengo
Published: (2021)
Adult Neurogenesis in the Human Dentate Gyrus
by: Fred H. Gage
Published: (2024)
by: Fred H. Gage
Published: (2024)
DROIDCCT: Cryptographic Compliance Test via Trillion-Scale Measurement
by: Moghimi, Daniel, et al.
Published: (2026)
by: Moghimi, Daniel, et al.
Published: (2026)
Profiling Resilient to Change in Probe Position
by: Bursztein, Elie, et al.
Published: (2026)
by: Bursztein, Elie, et al.
Published: (2026)
Generalized Power Attacks against Crypto Hardware using Long-Range Deep Learning
by: Bursztein, Elie, et al.
Published: (2023)
by: Bursztein, Elie, et al.
Published: (2023)
Program Booklet of the 4th African Student Council Symposium
by: Caivil Ndobela, et al.
Published: (2025)
by: Caivil Ndobela, et al.
Published: (2025)
The education Fellow's role in collaboratively implementing change
by: Ambika Wakhlu, et al.
Published: (2025)
by: Ambika Wakhlu, et al.
Published: (2025)
Three-body interactions in Rabi-coupled Bose gases: a perturbative approach
by: Tiengo, S, et al.
Published: (2025)
by: Tiengo, S, et al.
Published: (2025)
Frontier AI's Impact on the Cybersecurity Landscape
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial Attacks
by: Nasr, Milad, et al.
Published: (2025)
by: Nasr, Milad, et al.
Published: (2025)
Large-scale circulation with small diapycnal diffusion: The two-thermocline limit
by: Samelson, R., Vallis, G
Published: (1997)
by: Samelson, R., Vallis, G
Published: (1997)
Emergence of Fofonoff states in inviscid and viscous ocean circulation models
by: Wang, J., Vallis, G
Published: (1994)
by: Wang, J., Vallis, G
Published: (1994)
Give and Take: An End-To-End Investigation of Giveaway Scam Conversion Rates
by: Liu, Enze, et al.
Published: (2024)
by: Liu, Enze, et al.
Published: (2024)
Comparison of Scoring Rationales Between Large Language Models and Human Raters
by: Hua, Haowei, et al.
Published: (2025)
by: Hua, Haowei, et al.
Published: (2025)
Online Hate and Harmful Content
by: Keipi, Teo, et al.
Published: (2025)
by: Keipi, Teo, et al.
Published: (2025)
Harmful Suicide Content Detection
by: Park, Kyumin, et al.
Published: (2024)
by: Park, Kyumin, et al.
Published: (2024)
Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge
by: Wu, Junjie, et al.
Published: (2026)
by: Wu, Junjie, et al.
Published: (2026)
IRT Observed‐Score Equating for Rater‐Mediated Assessments Using a Hierarchical Rater Model
by: Tong Wu, et al.
Published: (2025)
by: Tong Wu, et al.
Published: (2025)
Rater Cohesion and Quality from a Vicarious Perspective
by: Pandita, Deepak, et al.
Published: (2024)
by: Pandita, Deepak, et al.
Published: (2024)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model
by: Jiao, Hong, et al.
Published: (2025)
by: Jiao, Hong, et al.
Published: (2025)
Mechanisms of Superrotation in Slowly-Rotating and Tidally-Locked Planets
by: Nicolas, Quentin, et al.
Published: (2025)
by: Nicolas, Quentin, et al.
Published: (2025)
Impact of COVID‐19 pandemic on sleep parameters and characteristics in individuals living with overweight and obesity
by: Stephen A. Glazer, et al.
Published: (2024)
by: Stephen A. Glazer, et al.
Published: (2024)
The Problem of Conserving Medieval Paintings on Exteriors in Slovenia
by: Šeme, Blaž
Published: (2007)
by: Šeme, Blaž
Published: (2007)
Social Media Bot Detection Research: Review of Literature
by: Rodič, Blaž
Published: (2025)
by: Rodič, Blaž
Published: (2025)
Similar Items
-
RETVec: Resilient and Efficient Text Vectorizer
by: Bursztein, Elie, et al.
Published: (2023) -
Understanding Help-Seeking and Help-Giving on Social Media for Image-Based Sexual Abuse
by: Wei, Miranda, et al.
Published: (2024) -
"It didn't feel right but I needed a job so desperately": Understanding People's Emotions & Help Needs During Financial Scams
by: Chanenson, Jake, et al.
Published: (2026) -
Understanding Help Seeking for Digital Privacy, Safety, and Security
by: Thomas, Kurt, et al.
Published: (2026) -
Magika: AI-Powered Content-Type Detection
by: Fratantonio, Yanick, et al.
Published: (2024)