BAID: A Benchmark for Bias Assessment of AI Detectors
Fuente:
arXiv
Saved in:
| Main Authors: | Basu, Priyam, Zhang, Yunfeng, Raheja, Vipul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
by: Gupta, Vipul, et al.
Published: (2023)
by: Gupta, Vipul, et al.
Published: (2023)
The Amazing Agent Race: Strong Tool Users, Weak Navigators
by: Kim, Zae Myung, et al.
Published: (2026)
by: Kim, Zae Myung, et al.
Published: (2026)
Benchmarking Cognitive Biases in Large Language Models as Evaluators
by: Koo, Ryan, et al.
Published: (2023)
by: Koo, Ryan, et al.
Published: (2023)
Sociodemographic Bias in Language Models: A Survey and Forward Path
by: Gupta, Vipul, et al.
Published: (2023)
by: Gupta, Vipul, et al.
Published: (2023)
TrueGradeAI: Retrieval-Augmented and Bias-Resistant AI for Transparent and Explainable Digital Assessments
by: Thakur, Rakesh, et al.
Published: (2025)
by: Thakur, Rakesh, et al.
Published: (2025)
Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
by: Mooney, James, et al.
Published: (2025)
by: Mooney, James, et al.
Published: (2025)
Explanations as Bias Detectors: A Critical Study of Local Post-hoc XAI Methods for Fairness Exploration
by: Papanikou, Vasiliki, et al.
Published: (2025)
by: Papanikou, Vasiliki, et al.
Published: (2025)
Knoop: Practical Enhancement of Knockoff with Over-Parameterization for Variable Selection
by: Zhang, Xiaochen, et al.
Published: (2025)
by: Zhang, Xiaochen, et al.
Published: (2025)
seeBias: A Comprehensive Tool for Assessing and Visualizing AI Fairness
by: Ning, Yilin, et al.
Published: (2025)
by: Ning, Yilin, et al.
Published: (2025)
Contextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models
by: Basu, Abhinaba, et al.
Published: (2026)
by: Basu, Abhinaba, et al.
Published: (2026)
Agentic Personalisation of Cross-Channel Marketing Experiences
by: Abboud, Sami, et al.
Published: (2025)
by: Abboud, Sami, et al.
Published: (2025)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
by: Gupta, Vipul, et al.
Published: (2024)
by: Gupta, Vipul, et al.
Published: (2024)
Multiplicative learning from observation-prediction ratios
by: Kim, Han, et al.
Published: (2025)
by: Kim, Han, et al.
Published: (2025)
Base Models Look Human To AI Detectors
by: Xu, Yixuan Even, et al.
Published: (2026)
by: Xu, Yixuan Even, et al.
Published: (2026)
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
by: Ranganath, Suraj, et al.
Published: (2026)
by: Ranganath, Suraj, et al.
Published: (2026)
Self-Attribution Bias: When AI Monitors Go Easy on Themselves
by: Khullar, Dipika, et al.
Published: (2026)
by: Khullar, Dipika, et al.
Published: (2026)
Enhancing AI-Based Tropical Cyclone Track and Intensity Forecasting via Systematic Bias Correction
by: Niu, Peisong, et al.
Published: (2026)
by: Niu, Peisong, et al.
Published: (2026)
Mathematics and Coding are Universal AI Benchmarks
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Assessing Trustworthiness of AI Training Dataset using Subjective Logic -- A Use Case on Bias
by: Ouattara, Koffi Ismael, et al.
Published: (2025)
by: Ouattara, Koffi Ismael, et al.
Published: (2025)
A Benchmark of Causal vs. Correlation AI for Predictive Maintenance
by: Dhande, Shaunak, et al.
Published: (2025)
by: Dhande, Shaunak, et al.
Published: (2025)
PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
by: Liu, Xinge, et al.
Published: (2026)
by: Liu, Xinge, et al.
Published: (2026)
Do Physics Foundation Models Learn Generalizable Physics? A Bias-Aware Benchmark Across Physical Regimes and Distribution Shifts
by: Chu, Mengdi, et al.
Published: (2026)
by: Chu, Mengdi, et al.
Published: (2026)
AuthorMist: Evading AI Text Detectors with Reinforcement Learning
by: David, Isaac, et al.
Published: (2025)
by: David, Isaac, et al.
Published: (2025)
Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI
by: Böck, Adrian Jaques, et al.
Published: (2024)
by: Böck, Adrian Jaques, et al.
Published: (2024)
CONCLAD: COntinuous Novel CLAss Detector
by: Rios, Amanda, et al.
Published: (2024)
by: Rios, Amanda, et al.
Published: (2024)
A Decision-driven Methodology for Designing Uncertainty-aware AI Self-Assessment
by: Canal, Gregory, et al.
Published: (2024)
by: Canal, Gregory, et al.
Published: (2024)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
by: Zhao, Kangwen, et al.
Published: (2025)
by: Zhao, Kangwen, et al.
Published: (2025)
Navigating the Shadows: Unveiling Effective Disturbances for Modern AI Content Detectors
by: Zhou, Ying, et al.
Published: (2024)
by: Zhou, Ying, et al.
Published: (2024)
Benchmarking is Broken -- Don't Let AI be its Own Judge
by: Cheng, Zerui, et al.
Published: (2025)
by: Cheng, Zerui, et al.
Published: (2025)
A Compression Perspective on Simplicity Bias
by: Marty, Tom, et al.
Published: (2026)
by: Marty, Tom, et al.
Published: (2026)
Learning Discriminative and Generalizable Anomaly Detector for Dynamic Graph with Limited Supervision
by: Tian, Yuxing, et al.
Published: (2026)
by: Tian, Yuxing, et al.
Published: (2026)
A Theoretical Understanding of Gradient Bias in Meta-Reinforcement Learning
by: Feng, Xidong, et al.
Published: (2021)
by: Feng, Xidong, et al.
Published: (2021)
EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
by: Block, Adam, et al.
Published: (2025)
by: Block, Adam, et al.
Published: (2025)
AI for Water Sustainability: Global Water Quality Assessment and Prediction with Explainable AI with LLM Chatbot for Insights
by: Paneru, Biplov, et al.
Published: (2024)
by: Paneru, Biplov, et al.
Published: (2024)
IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text Detector
by: Chen, Zheng, et al.
Published: (2025)
by: Chen, Zheng, et al.
Published: (2025)
PepBenchmark: A Standardized Benchmark for Peptide Machine Learning
by: Zhang, Jiahui, et al.
Published: (2026)
by: Zhang, Jiahui, et al.
Published: (2026)
Generative AI for Health Technology Assessment: Opportunities, Challenges, and Policy Considerations
by: Fleurence, Rachael, et al.
Published: (2024)
by: Fleurence, Rachael, et al.
Published: (2024)
HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI
by: Pricope, Tidor-Vlad
Published: (2025)
by: Pricope, Tidor-Vlad
Published: (2025)
BEExAI: Benchmark to Evaluate Explainable AI
by: Sithakoul, Samuel, et al.
Published: (2024)
by: Sithakoul, Samuel, et al.
Published: (2024)
Similar Items
-
CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
by: Gupta, Vipul, et al.
Published: (2023) -
The Amazing Agent Race: Strong Tool Users, Weak Navigators
by: Kim, Zae Myung, et al.
Published: (2026) -
Benchmarking Cognitive Biases in Large Language Models as Evaluators
by: Koo, Ryan, et al.
Published: (2023) -
Sociodemographic Bias in Language Models: A Survey and Forward Path
by: Gupta, Vipul, et al.
Published: (2023) -
TrueGradeAI: Retrieval-Augmented and Bias-Resistant AI for Transparent and Explainable Digital Assessments
by: Thakur, Rakesh, et al.
Published: (2025)