CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Vipul, Venkit, Pranav Narayanan, Laurençon, Hugo, Wilson, Shomir, Passonneau, Rebecca J. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sociodemographic Bias in Language Models: A Survey and Forward Path
by: Gupta, Vipul, et al.
Published: (2023)
by: Gupta, Vipul, et al.
Published: (2023)
An Audit on the Perspectives and Challenges of Hallucinations in NLP
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
by: Gupta, Vipul, et al.
Published: (2024)
by: Gupta, Vipul, et al.
Published: (2024)
Do Generative AI Models Output Harm while Representing Non-Western Cultures: Evidence from A Community-Centered Approach
by: Ghosh, Sourojit, et al.
Published: (2024)
by: Ghosh, Sourojit, et al.
Published: (2024)
NLP Meets the World: Toward Improving Conversations With the Public About Natural Language Processing Research
by: Wilson, Shomir
Published: (2025)
by: Wilson, Shomir
Published: (2025)
VerAs: Verify then Assess STEM Lab Reports
by: Atil, Berk, et al.
Published: (2024)
by: Atil, Berk, et al.
Published: (2024)
Model Unlearning Objectives Vary for Distinct Language Functions
by: Atil, Berk, et al.
Published: (2026)
by: Atil, Berk, et al.
Published: (2026)
InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
The Need for a Socially-Grounded Persona Framework for User Simulation
by: Venkit, Pranav Narayanan, et al.
Published: (2026)
by: Venkit, Pranav Narayanan, et al.
Published: (2026)
From Melting Pots to Misrepresentations: Exploring Harms in Generative AI
by: Gautam, Sanjana, et al.
Published: (2024)
by: Gautam, Sanjana, et al.
Published: (2024)
BAID: A Benchmark for Bias Assessment of AI Detectors
by: Basu, Priyam, et al.
Published: (2025)
by: Basu, Priyam, et al.
Published: (2025)
Learning to Trust the Crowd: A Multi-Model Consensus Reasoning Engine for Large Language Models
by: Kallem, Pranav
Published: (2026)
by: Kallem, Pranav
Published: (2026)
Hey GPT, Can You be More Racist? Analysis from Crowdsourced Attempts to Elicit Biased Content from Generative AI
by: Guo, Hangzhi, et al.
Published: (2024)
by: Guo, Hangzhi, et al.
Published: (2024)
CALM: Curiosity-Driven Auditing for Large Language Models
by: Zheng, Xiang, et al.
Published: (2025)
by: Zheng, Xiang, et al.
Published: (2025)
Something Just Like TRuST : Toxicity Recognition of Span and Target
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
by: Zhao, Qihao, et al.
Published: (2024)
by: Zhao, Qihao, et al.
Published: (2024)
ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment
by: Naous, Tarek, et al.
Published: (2023)
by: Naous, Tarek, et al.
Published: (2023)
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
On-device System of Compositional Multi-tasking in Large Language Models
by: Bohdal, Ondrej, et al.
Published: (2025)
by: Bohdal, Ondrej, et al.
Published: (2025)
Efficient Compositional Multi-tasking for On-device Large Language Models
by: Bohdal, Ondrej, et al.
Published: (2025)
by: Bohdal, Ondrej, et al.
Published: (2025)
Race and Privacy in Broadcast Police Communications
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
by: Tang, Zhengyang, et al.
Published: (2025)
by: Tang, Zhengyang, et al.
Published: (2025)
Benchmarking Cognitive Biases in Large Language Models as Evaluators
by: Koo, Ryan, et al.
Published: (2023)
by: Koo, Ryan, et al.
Published: (2023)
Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
by: Park, Jean, et al.
Published: (2024)
by: Park, Jean, et al.
Published: (2024)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Benchmarking Gender and Political Bias in Large Language Models
by: Yang, Jinrui, et al.
Published: (2025)
by: Yang, Jinrui, et al.
Published: (2025)
CTBench: A Comprehensive Benchmark for Evaluating Language Model Capabilities in Clinical Trial Design
by: Neehal, Nafis, et al.
Published: (2024)
by: Neehal, Nafis, et al.
Published: (2024)
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery
by: Song, Xiaoshuai, et al.
Published: (2024)
by: Song, Xiaoshuai, et al.
Published: (2024)
Can Third-parties Read Our Emotions?
by: Li, Jiayi, et al.
Published: (2025)
by: Li, Jiayi, et al.
Published: (2025)
An Electrocardiogram Multi-task Benchmark with Comprehensive Evaluations and Insightful Findings
by: Xu, Yuhao, et al.
Published: (2025)
by: Xu, Yuhao, et al.
Published: (2025)
Understanding and Mitigating Tokenization Bias in Language Models
by: Phan, Buu, et al.
Published: (2024)
by: Phan, Buu, et al.
Published: (2024)
Aviary: training language agents on challenging scientific tasks
by: Narayanan, Siddharth, et al.
Published: (2024)
by: Narayanan, Siddharth, et al.
Published: (2024)
Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models
by: Islam, Mohammed Saidul, et al.
Published: (2026)
by: Islam, Mohammed Saidul, et al.
Published: (2026)
RadLite: Multi-Task LoRA Fine-Tuning of Small Language Models for CPU-Deployable Radiology AI
by: Gupta, Pankaj, et al.
Published: (2026)
by: Gupta, Pankaj, et al.
Published: (2026)
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
Benchmarking Benchmark Leakage in Large Language Models
by: Xu, Ruijie, et al.
Published: (2024)
by: Xu, Ruijie, et al.
Published: (2024)
The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models
by: Jeong, Daniel P., et al.
Published: (2024)
by: Jeong, Daniel P., et al.
Published: (2024)
Similar Items
-
Sociodemographic Bias in Language Models: A Survey and Forward Path
by: Gupta, Vipul, et al.
Published: (2023) -
An Audit on the Perspectives and Challenges of Hallucinations in NLP
by: Venkit, Pranav Narayanan, et al.
Published: (2024) -
A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
by: Venkit, Pranav Narayanan, et al.
Published: (2025) -
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
by: Gupta, Vipul, et al.
Published: (2024) -
Do Generative AI Models Output Harm while Representing Non-Western Cultures: Evidence from A Community-Centered Approach
by: Ghosh, Sourojit, et al.
Published: (2024)