Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree
Fuente:
arXiv
Saved in:
| Main Authors: | Jaggi, Harbani, Murali, Kashyap, Fleisig, Eve, Bıyık, Erdem |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
by: Fleisig, Eve, et al.
Published: (2023)
by: Fleisig, Eve, et al.
Published: (2023)
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
by: Fleisig, Eve, et al.
Published: (2025)
by: Fleisig, Eve, et al.
Published: (2025)
First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models
by: Saphra, Naomi, et al.
Published: (2023)
by: Saphra, Naomi, et al.
Published: (2023)
Will Annotators Disagree? Identifying Subjectivity in Value-Laden Arguments
by: Homayounirad, Amir, et al.
Published: (2025)
by: Homayounirad, Amir, et al.
Published: (2025)
Ghostbuster: Detecting Text Ghostwritten by Large Language Models
by: Verma, Vivek, et al.
Published: (2023)
by: Verma, Vivek, et al.
Published: (2023)
Standard Language Ideology in AI-Generated Language
by: Smith, Genevieve, et al.
Published: (2024)
by: Smith, Genevieve, et al.
Published: (2024)
The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
by: Fleisig, Eve, et al.
Published: (2024)
by: Fleisig, Eve, et al.
Published: (2024)
Diverging Preferences: When do Annotators Disagree and do Models Know?
by: Zhang, Michael JQ, et al.
Published: (2024)
by: Zhang, Michael JQ, et al.
Published: (2024)
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
by: Fleisig, Eve, et al.
Published: (2024)
by: Fleisig, Eve, et al.
Published: (2024)
When Annotators Agree but Labels Disagree: The Projection Problem in Stance Detection
by: Zhang, Bowen
Published: (2026)
by: Zhang, Bowen
Published: (2026)
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
by: Sung, Yoo Yeon, et al.
Published: (2024)
by: Sung, Yoo Yeon, et al.
Published: (2024)
GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration
by: Sung, Yoo Yeon, et al.
Published: (2025)
by: Sung, Yoo Yeon, et al.
Published: (2025)
Learning Who Disagrees: Demographic Importance Weighting for Modeling Annotator Distributions with DiADEM
by: Shetty, Samay U., et al.
Published: (2026)
by: Shetty, Samay U., et al.
Published: (2026)
Mapping Social Choice Theory to RLHF
by: Dai, Jessica, et al.
Published: (2024)
by: Dai, Jessica, et al.
Published: (2024)
When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and Ambiguity
by: Rair, Nisrine, et al.
Published: (2025)
by: Rair, Nisrine, et al.
Published: (2025)
PSK@EEUCA 2026: Fine-Tuning Large Language Models with Synthetic Data Augmentation for Multi-Class Toxicity Detection in Gaming Chat
by: Pulipaka, Srikar Kashyap
Published: (2026)
by: Pulipaka, Srikar Kashyap
Published: (2026)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
by: Subramonian, Arjun, et al.
Published: (2025)
by: Subramonian, Arjun, et al.
Published: (2025)
AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
by: Gong, Litian, et al.
Published: (2025)
by: Gong, Litian, et al.
Published: (2025)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
by: Gor, Maharshi, et al.
Published: (2026)
by: Gor, Maharshi, et al.
Published: (2026)
ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality
by: Luo, Yu-Xiang, et al.
Published: (2025)
by: Luo, Yu-Xiang, et al.
Published: (2025)
Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains
by: Schmidt, Finn, et al.
Published: (2026)
by: Schmidt, Finn, et al.
Published: (2026)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
by: Muscato, Benedetta, et al.
Published: (2026)
by: Muscato, Benedetta, et al.
Published: (2026)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)
by: Messmer, Bettina, et al.
Published: (2025)
Turkish Delights: a Dataset on Turkish Euphemisms
by: Biyik, Hasan Can, et al.
Published: (2024)
by: Biyik, Hasan Can, et al.
Published: (2024)
Training robots with natural and lightweight human feedback
by: Erdem Bıyık
Published: (2026)
by: Erdem Bıyık
Published: (2026)
Finding Pareto Trade-offs in Fair and Accurate Detection of Toxic Speech
by: Gupta, Soumyajit, et al.
Published: (2022)
by: Gupta, Soumyajit, et al.
Published: (2022)
Mitigating Biases to Embrace Diversity: A Comprehensive Annotation Benchmark for Toxic Language
by: Hou, Xinmeng
Published: (2024)
by: Hou, Xinmeng
Published: (2024)
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
by: Hong, Pingjun, et al.
Published: (2025)
by: Hong, Pingjun, et al.
Published: (2025)
Legal Experts Disagree With Rationale Extraction Techniques for Explaining ECtHR Case Outcome Classification
by: Namazov, Mahammad, et al.
Published: (2026)
by: Namazov, Mahammad, et al.
Published: (2026)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
by: Turki, Yassine, et al.
Published: (2026)
by: Turki, Yassine, et al.
Published: (2026)
Diagnosing Hate Speech Classification: Where Do Humans and Machines Disagree, and Why?
by: Yang, Xilin
Published: (2024)
by: Yang, Xilin
Published: (2024)
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
Efficient Learning Content Retrieval with Knowledge Injection
by: Sariturk, Batuhan, et al.
Published: (2024)
by: Sariturk, Batuhan, et al.
Published: (2024)
When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews
by: Kumar, Sandeep, et al.
Published: (2026)
by: Kumar, Sandeep, et al.
Published: (2026)
Core-based Hierarchies for Efficient GraphRAG
by: Hossain, Jakir, et al.
Published: (2026)
by: Hossain, Jakir, et al.
Published: (2026)
Why They Disagree: Decoding Differences in Opinions about AI Risk on the Lex Fridman Podcast
by: Truong, Nghi, et al.
Published: (2025)
by: Truong, Nghi, et al.
Published: (2025)
When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
iTAG: Inverse Design for Natural Text Generation with Accurate Causal Graph Annotations
by: Wang, Wenshuo, et al.
Published: (2026)
by: Wang, Wenshuo, et al.
Published: (2026)
Similar Items
-
When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
by: Fleisig, Eve, et al.
Published: (2023) -
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
by: Fleisig, Eve, et al.
Published: (2025) -
First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models
by: Saphra, Naomi, et al.
Published: (2023) -
Will Annotators Disagree? Identifying Subjectivity in Value-Laden Arguments
by: Homayounirad, Amir, et al.
Published: (2025) -
Ghostbuster: Detecting Text Ghostwritten by Large Language Models
by: Verma, Vivek, et al.
Published: (2023)