Fair and Calibrated Toxicity Detection with Robust Training and Abstention
Fuente:
arXiv
Saved in:
| Main Author: | Surana, Mokshit |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study
by: Surana, Mokshit, et al.
Published: (2026)
by: Surana, Mokshit, et al.
Published: (2026)
Omni-Modal Dissonance Benchmark: Systematically Breaking Modality Consensus to Probe Robustness and Calibrated Abstention
by: Nazi, Zabir Al, et al.
Published: (2026)
by: Nazi, Zabir Al, et al.
Published: (2026)
Geometry-Calibrated Conformal Abstention for Language Models
by: Xu, Rui, et al.
Published: (2026)
by: Xu, Rui, et al.
Published: (2026)
BOKBO (Best of K Bad Options): Calibrated Abstention for VLA Policies
by: Singh, Anya, et al.
Published: (2026)
by: Singh, Anya, et al.
Published: (2026)
FairJudge: Abstention-Aware Multimodal Judges for Fairness and Alignment Evaluation in Text-to-Image Models
by: Sahili, Zahraa Al, et al.
Published: (2025)
by: Sahili, Zahraa Al, et al.
Published: (2025)
Policy Learning with Abstention
by: Sawarni, Ayush, et al.
Published: (2025)
by: Sawarni, Ayush, et al.
Published: (2025)
Efficient Active Learning with Abstention
by: Zhu, Yinglun, et al.
Published: (2022)
by: Zhu, Yinglun, et al.
Published: (2022)
Bandits with Abstention under Expert Advice
by: Pasteris, Stephen, et al.
Published: (2024)
by: Pasteris, Stephen, et al.
Published: (2024)
Structured Prediction with Abstention via the Lovász Hinge
by: Finocchiaro, Jessie, et al.
Published: (2025)
by: Finocchiaro, Jessie, et al.
Published: (2025)
Bounded-Abstention Multi-horizon Time-series Forecasting
by: Stradiotti, Luca, et al.
Published: (2026)
by: Stradiotti, Luca, et al.
Published: (2026)
When In Doubt, Abstain: The Impact of Abstention on Strategic Classification
by: Alkarmi, Lina, et al.
Published: (2025)
by: Alkarmi, Lina, et al.
Published: (2025)
Distribution-Free Sequential Prediction with Abstentions
by: Yu, Jialin, et al.
Published: (2026)
by: Yu, Jialin, et al.
Published: (2026)
Generative Inverse Design with Abstention via Diagonal Flow Matching
by: de Campos, Miguel, et al.
Published: (2026)
by: de Campos, Miguel, et al.
Published: (2026)
Predictor-Rejector Multi-Class Abstention: Theoretical Analysis and Algorithms
by: Mao, Anqi, et al.
Published: (2023)
by: Mao, Anqi, et al.
Published: (2023)
Effective Skill Unlearning through Intervention and Abstention
by: Li, Yongce, et al.
Published: (2025)
by: Li, Yongce, et al.
Published: (2025)
Bounded-Abstention Pairwise Learning to Rank
by: Ferrara, Antonio, et al.
Published: (2025)
by: Ferrara, Antonio, et al.
Published: (2025)
Theory and Algorithms for Learning with Multi-Class Abstention and Multi-Expert Deferral
by: Mao, Anqi
Published: (2025)
by: Mao, Anqi
Published: (2025)
Online Conformal Abstention for Factuality Control Under Adversarial Bandit Feedback
by: Lee, Minjae, et al.
Published: (2025)
by: Lee, Minjae, et al.
Published: (2025)
Adversarial Resilience in Sequential Prediction via Abstention
by: Goel, Surbhi, et al.
Published: (2023)
by: Goel, Surbhi, et al.
Published: (2023)
MKA: Leveraging Cross-Lingual Consensus for Model Abstention
by: Duwal, Sharad
Published: (2025)
by: Duwal, Sharad
Published: (2025)
Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs
by: Varga-Umbrich, Eszter, et al.
Published: (2026)
by: Varga-Umbrich, Eszter, et al.
Published: (2026)
Threshold-Independent Fair Matching through Score Calibration
by: Moslemi, Mohammad Hossein, et al.
Published: (2024)
by: Moslemi, Mohammad Hossein, et al.
Published: (2024)
Theoretically Grounded Loss Functions and Algorithms for Score-Based Multi-Class Abstention
by: Mao, Anqi, et al.
Published: (2023)
by: Mao, Anqi, et al.
Published: (2023)
Mitigating LLM Hallucinations via Conformal Abstention
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
Fairness and Robustness in Machine Unlearning
by: Tran, Khoa, et al.
Published: (2025)
by: Tran, Khoa, et al.
Published: (2025)
An Effective, Robust and Fairness-aware Hate Speech Detection Framework
by: Mou, Guanyi, et al.
Published: (2024)
by: Mou, Guanyi, et al.
Published: (2024)
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
by: Hedström, Anna, et al.
Published: (2025)
by: Hedström, Anna, et al.
Published: (2025)
Projected Boosting with Fairness Constraints: Quantifying the Cost of Fair Training Distributions
by: Asiaee, Amir, et al.
Published: (2026)
by: Asiaee, Amir, et al.
Published: (2026)
On the Role of Calibration in Benchmarking Algorithmic Fairness for Skin Cancer Detection
by: Dominique, Brandon, et al.
Published: (2025)
by: Dominique, Brandon, et al.
Published: (2025)
Reliable Abstention under Adversarial Injections: Tight Lower Bounds and New Upper Bounds
by: Edelman, Ezra, et al.
Published: (2026)
by: Edelman, Ezra, et al.
Published: (2026)
Calibrated Credit Intelligence: Shift-Robust and Fair Risk Scoring with Bayesian Uncertainty and Gradient Boosting
by: Nayak, Srikumar
Published: (2026)
by: Nayak, Srikumar
Published: (2026)
Fairness-aware Anomaly Detection via Fair Projection
by: Xiao, Feng, et al.
Published: (2025)
by: Xiao, Feng, et al.
Published: (2025)
Detecting Toxic Flow
by: Cartea, Álvaro, et al.
Published: (2023)
by: Cartea, Álvaro, et al.
Published: (2023)
Learning When Not to Learn: Risk-Sensitive Abstention in Bandits with Unbounded Rewards
by: Liaw, Sarah, et al.
Published: (2025)
by: Liaw, Sarah, et al.
Published: (2025)
Asymptotically and Minimax Optimal Regret Bounds for Multi-Armed Bandits with Abstention
by: Yang, Junwen, et al.
Published: (2024)
by: Yang, Junwen, et al.
Published: (2024)
How Robust is your Fair Model? Exploring the Robustness of Diverse Fairness Strategies
by: Small, Edward, et al.
Published: (2022)
by: Small, Edward, et al.
Published: (2022)
Predict Confidently, Predict Right: Abstention in Dynamic Graph Learning
by: Gayen, Jayadratha, et al.
Published: (2025)
by: Gayen, Jayadratha, et al.
Published: (2025)
Beyond Confidence: Adaptive Abstention in Dual-Threshold Conformal Prediction for Autonomous System Perception
by: Kumar, Divake, et al.
Published: (2025)
by: Kumar, Divake, et al.
Published: (2025)
Fairness in Survival Analysis with Distributionally Robust Optimization
by: Hu, Shu, et al.
Published: (2024)
by: Hu, Shu, et al.
Published: (2024)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
by: Nakkiran, Preetum, et al.
Published: (2025)
by: Nakkiran, Preetum, et al.
Published: (2025)
Similar Items
-
Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study
by: Surana, Mokshit, et al.
Published: (2026) -
Omni-Modal Dissonance Benchmark: Systematically Breaking Modality Consensus to Probe Robustness and Calibrated Abstention
by: Nazi, Zabir Al, et al.
Published: (2026) -
Geometry-Calibrated Conformal Abstention for Language Models
by: Xu, Rui, et al.
Published: (2026) -
BOKBO (Best of K Bad Options): Calibrated Abstention for VLA Policies
by: Singh, Anya, et al.
Published: (2026) -
FairJudge: Abstention-Aware Multimodal Judges for Fairness and Alignment Evaluation in Text-to-Image Models
by: Sahili, Zahraa Al, et al.
Published: (2025)