COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Govil, Priyanshul, Jain, Hemang, Bonagiri, Vamshi Krishna, Chadha, Aman, Kumaraguru, Ponnurangam, Gaur, Manas, Dey, Sanorita |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SaGE: Evaluating Moral Consistency in Large Language Models
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024)
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024)
Measuring Moral Inconsistencies in Large Language Models
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024)
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024)
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
di: Aneja, Krishak, et al.
Pubblicazione: (2026)
di: Aneja, Krishak, et al.
Pubblicazione: (2026)
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
di: Kodali, Prashant, et al.
Pubblicazione: (2024)
di: Kodali, Prashant, et al.
Pubblicazione: (2024)
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
di: Joishy, Anish R, et al.
Pubblicazione: (2025)
di: Joishy, Anish R, et al.
Pubblicazione: (2025)
K-PERM: Personalized Response Generation Using Dynamic Knowledge Retrieval and Persona-Adaptive Queries
di: Raj, Kanak, et al.
Pubblicazione: (2023)
di: Raj, Kanak, et al.
Pubblicazione: (2023)
Just KIDDIN: Knowledge Infusion and Distillation for Detection of INdecent Memes
di: Garg, Rahul, et al.
Pubblicazione: (2024)
di: Garg, Rahul, et al.
Pubblicazione: (2024)
KnowledgePrompts: Exploring the Abilities of Large Language Models to Solve Proportional Analogies via Knowledge-Enhanced Prompting
di: Wijesiriwardene, Thilini, et al.
Pubblicazione: (2024)
di: Wijesiriwardene, Thilini, et al.
Pubblicazione: (2024)
Unboxing Occupational Bias: Grounded Debiasing of LLMs with U.S. Labor Data
di: Gorti, Atmika, et al.
Pubblicazione: (2024)
di: Gorti, Atmika, et al.
Pubblicazione: (2024)
TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
di: Kavathekar, Ishan, et al.
Pubblicazione: (2025)
di: Kavathekar, Ishan, et al.
Pubblicazione: (2025)
Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2025)
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2025)
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
di: Vedula, Bhaskara Hanuma, et al.
Pubblicazione: (2026)
di: Vedula, Bhaskara Hanuma, et al.
Pubblicazione: (2026)
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
di: Das, Nilanjana, et al.
Pubblicazione: (2024)
di: Das, Nilanjana, et al.
Pubblicazione: (2024)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
di: Haider, Batool, et al.
Pubblicazione: (2025)
di: Haider, Batool, et al.
Pubblicazione: (2025)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
di: Krishnappa, Pushwitha, et al.
Pubblicazione: (2026)
di: Krishnappa, Pushwitha, et al.
Pubblicazione: (2026)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
di: Singh, Smriti, et al.
Pubblicazione: (2024)
di: Singh, Smriti, et al.
Pubblicazione: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2025)
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2025)
Multilingual Coreference Resolution in Low-resource South Asian Languages
di: Mishra, Ritwik, et al.
Pubblicazione: (2024)
di: Mishra, Ritwik, et al.
Pubblicazione: (2024)
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
di: KJ, Sankalp, et al.
Pubblicazione: (2025)
di: KJ, Sankalp, et al.
Pubblicazione: (2025)
PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI
di: Mukhopadhyay, Srija, et al.
Pubblicazione: (2025)
di: Mukhopadhyay, Srija, et al.
Pubblicazione: (2025)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
di: Khoshnoodi, Mahsa, et al.
Pubblicazione: (2024)
di: Khoshnoodi, Mahsa, et al.
Pubblicazione: (2024)
Multilingual State Space Models for Structured Question Answering in Indic Languages
di: Vats, Arpita, et al.
Pubblicazione: (2025)
di: Vats, Arpita, et al.
Pubblicazione: (2025)
Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability
di: Aggarwal, Yash, et al.
Pubblicazione: (2026)
di: Aggarwal, Yash, et al.
Pubblicazione: (2026)
Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2026)
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2026)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
di: Saha, Anusa, et al.
Pubblicazione: (2026)
di: Saha, Anusa, et al.
Pubblicazione: (2026)
Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
di: Khatri, Mann, et al.
Pubblicazione: (2025)
di: Khatri, Mann, et al.
Pubblicazione: (2025)
How Culturally Aware are Vision-Language Models?
di: Burda-Lassen, Olena, et al.
Pubblicazione: (2024)
di: Burda-Lassen, Olena, et al.
Pubblicazione: (2024)
Manifold-based Sampling for In-Context Hallucination Detection in Large Language Models
di: Vamshi, Bodla Krishna, et al.
Pubblicazione: (2026)
di: Vamshi, Bodla Krishna, et al.
Pubblicazione: (2026)
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
di: Ghosh, Akash, et al.
Pubblicazione: (2024)
di: Ghosh, Akash, et al.
Pubblicazione: (2024)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
di: Das, Nilanjana, et al.
Pubblicazione: (2026)
di: Das, Nilanjana, et al.
Pubblicazione: (2026)
Long-context Non-factoid Question Answering in Indic Languages
di: Mishra, Ritwik, et al.
Pubblicazione: (2025)
di: Mishra, Ritwik, et al.
Pubblicazione: (2025)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
Multilingual Non-Factoid Question Answering with Answer Paragraph Selection
di: Mishra, Ritwik, et al.
Pubblicazione: (2024)
di: Mishra, Ritwik, et al.
Pubblicazione: (2024)
Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
SEMMA: A Semantic Aware Knowledge Graph Foundation Model
di: Arun, Arvindh, et al.
Pubblicazione: (2025)
di: Arun, Arvindh, et al.
Pubblicazione: (2025)
I Can't Believe It's Corrupt: Evaluating Corruption in Multi-Agent Governance Systems
di: P, Vedanta S, et al.
Pubblicazione: (2026)
di: P, Vedanta S, et al.
Pubblicazione: (2026)
On the Relationship between Sentence Analogy Identification and Sentence Structure Encoding in Large Language Models
di: Wijesiriwardene, Thilini, et al.
Pubblicazione: (2023)
di: Wijesiriwardene, Thilini, et al.
Pubblicazione: (2023)
Documenti analoghi
-
SaGE: Evaluating Moral Consistency in Large Language Models
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024) -
Measuring Moral Inconsistencies in Large Language Models
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024) -
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
di: Aneja, Krishak, et al.
Pubblicazione: (2026) -
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
di: Kodali, Prashant, et al.
Pubblicazione: (2024) -
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
di: Joishy, Anish R, et al.
Pubblicazione: (2025)