Counterfactual Probing for the Influence of Affect and Specificity on Intergroup Bias
Fuente:
arXiv
Salvato in:
| Autori principali: | Govindarajan, Venkata S, Mahowald, Kyle, Beaver, David I., Li, Junyi Jessy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Do they mean 'us'? Interpreting Referring Expressions in Intergroup Bias
di: Govindarajan, Venkata S, et al.
Pubblicazione: (2024)
di: Govindarajan, Venkata S, et al.
Pubblicazione: (2024)
How people talk about each other: Modeling Generalized Intergroup Bias and Emotion
di: Govindarajan, Venkata S, et al.
Pubblicazione: (2022)
di: Govindarajan, Venkata S, et al.
Pubblicazione: (2022)
Lil-Bevo: Explorations of Strategies for Training Language Models in More Humanlike Ways
di: Govindarajan, Venkata S, et al.
Pubblicazione: (2023)
di: Govindarajan, Venkata S, et al.
Pubblicazione: (2023)
Strategic Dialogue Assessment: The Crooked Path to Innocence
di: Zheng, Anshun Asher, et al.
Pubblicazione: (2025)
di: Zheng, Anshun Asher, et al.
Pubblicazione: (2025)
Help! Need Advice on Identifying Advice
di: Govindarajan, Venkata Subrahmanyan, et al.
Pubblicazione: (2020)
di: Govindarajan, Venkata Subrahmanyan, et al.
Pubblicazione: (2020)
Multimodal QUD: Inquisitive Questions from Scientific Figures
di: Wu, Yating, et al.
Pubblicazione: (2026)
di: Wu, Yating, et al.
Pubblicazione: (2026)
Template-Based Probes Are Imperfect Lenses for Counterfactual Bias Evaluation in LLMs
di: Kohankhaki, Farnaz, et al.
Pubblicazione: (2024)
di: Kohankhaki, Farnaz, et al.
Pubblicazione: (2024)
Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs
di: Sheffield, William, et al.
Pubblicazione: (2025)
di: Sheffield, William, et al.
Pubblicazione: (2025)
When AI Speaks, Whose Values Does It Express? A Cross-Cultural Audit of Individualism-Collectivism Bias in Large Language Models
di: Venkata, Pruthvinath Jeripity
Pubblicazione: (2026)
di: Venkata, Pruthvinath Jeripity
Pubblicazione: (2026)
Counterfactual LLM-based Framework for Measuring Rhetorical Style
di: Qiu, Jingyi, et al.
Pubblicazione: (2025)
di: Qiu, Jingyi, et al.
Pubblicazione: (2025)
Dark & Stormy: Modeling Humor in Sentences from the Bulwer-Lytton Fiction Contest
di: Govindarajan, Venkata S, et al.
Pubblicazione: (2025)
di: Govindarajan, Venkata S, et al.
Pubblicazione: (2025)
For Generated Text, Is NLI-Neutral Text the Best Text?
di: Mersinias, Michail, et al.
Pubblicazione: (2023)
di: Mersinias, Michail, et al.
Pubblicazione: (2023)
You Can't Fight in Here! This is BBS!
di: Futrell, Richard, et al.
Pubblicazione: (2026)
di: Futrell, Richard, et al.
Pubblicazione: (2026)
Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs
di: Misra, Kanishka, et al.
Pubblicazione: (2024)
di: Misra, Kanishka, et al.
Pubblicazione: (2024)
Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs
di: Boguraev, Sasha, et al.
Pubblicazione: (2026)
di: Boguraev, Sasha, et al.
Pubblicazione: (2026)
How Linguistics Learned to Stop Worrying and Love the Language Models
di: Futrell, Richard, et al.
Pubblicazione: (2025)
di: Futrell, Richard, et al.
Pubblicazione: (2025)
Are Language Models More Like Libraries or Like Librarians? Bibliotechnism, the Novel Reference Problem, and the Attitudes of LLMs
di: Lederman, Harvey, et al.
Pubblicazione: (2024)
di: Lederman, Harvey, et al.
Pubblicazione: (2024)
Rethinking LLM Bias Probing Using Lessons from the Social Sciences
di: Morehouse, Kirsten N., et al.
Pubblicazione: (2025)
di: Morehouse, Kirsten N., et al.
Pubblicazione: (2025)
Bridging or Breaking: Impact of Intergroup Interactions on Religious Polarization
di: Chaturvedi, Rochana, et al.
Pubblicazione: (2024)
di: Chaturvedi, Rochana, et al.
Pubblicazione: (2024)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
di: Ma, Mingyu Derek, et al.
Pubblicazione: (2023)
di: Ma, Mingyu Derek, et al.
Pubblicazione: (2023)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
di: Nemitz, Jonathan, et al.
Pubblicazione: (2026)
di: Nemitz, Jonathan, et al.
Pubblicazione: (2026)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
di: Yilmaz, Berk, et al.
Pubblicazione: (2025)
di: Yilmaz, Berk, et al.
Pubblicazione: (2025)
Unsupervised Concept Vector Extraction for Bias Control in LLMs
di: Cyberey, Hannah, et al.
Pubblicazione: (2025)
di: Cyberey, Hannah, et al.
Pubblicazione: (2025)
AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment
di: Li, Kun, et al.
Pubblicazione: (2025)
di: Li, Kun, et al.
Pubblicazione: (2025)
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
di: Cyberey, Hannah, et al.
Pubblicazione: (2024)
di: Cyberey, Hannah, et al.
Pubblicazione: (2024)
SWAY: A Counterfactual Computational Linguistic Approach to Measuring and Mitigating Sycophancy
di: Bhalla, Joy, et al.
Pubblicazione: (2026)
di: Bhalla, Joy, et al.
Pubblicazione: (2026)
Extracting Affect Aggregates from Longitudinal Social Media Data with Temporal Adapters for Large Language Models
di: Ahnert, Georg, et al.
Pubblicazione: (2024)
di: Ahnert, Georg, et al.
Pubblicazione: (2024)
Emergent Introspection in AI is Content-Agnostic
di: Lederman, Harvey, et al.
Pubblicazione: (2026)
di: Lederman, Harvey, et al.
Pubblicazione: (2026)
The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Models
di: Drucker, Daniel, et al.
Pubblicazione: (2026)
di: Drucker, Daniel, et al.
Pubblicazione: (2026)
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
di: Mo, Kaijie, et al.
Pubblicazione: (2026)
di: Mo, Kaijie, et al.
Pubblicazione: (2026)
Counterspeech for Mitigating the Influence of Media Bias: Comparing Human and LLM-Generated Responses
di: Lin, Luyang, et al.
Pubblicazione: (2025)
di: Lin, Luyang, et al.
Pubblicazione: (2025)
Trust, Safety, and Accuracy: Assessing LLMs for Routine Maternity Advice
di: Divya, V Sai, et al.
Pubblicazione: (2026)
di: Divya, V Sai, et al.
Pubblicazione: (2026)
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education
di: Liu, Naiming, et al.
Pubblicazione: (2024)
di: Liu, Naiming, et al.
Pubblicazione: (2024)
Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language Models
di: Faulborn, Mats, et al.
Pubblicazione: (2025)
di: Faulborn, Mats, et al.
Pubblicazione: (2025)
Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting
di: Kambhatla, Gauri, et al.
Pubblicazione: (2025)
di: Kambhatla, Gauri, et al.
Pubblicazione: (2025)
A Practical Method for Generating String Counterfactuals
di: Avitan, Matan, et al.
Pubblicazione: (2024)
di: Avitan, Matan, et al.
Pubblicazione: (2024)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
di: Ghosh, Rajarshi, et al.
Pubblicazione: (2025)
di: Ghosh, Rajarshi, et al.
Pubblicazione: (2025)
On The Conceptualization and Societal Impact of Cross-Cultural Bias
di: Bhandari, Vitthal
Pubblicazione: (2025)
di: Bhandari, Vitthal
Pubblicazione: (2025)
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
di: Sen, Indira, et al.
Pubblicazione: (2023)
di: Sen, Indira, et al.
Pubblicazione: (2023)
Participle-Prepended Nominals Have Lower Entropy Than Nominals Appended After the Participle
di: Denlinger, Kristie, et al.
Pubblicazione: (2024)
di: Denlinger, Kristie, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Do they mean 'us'? Interpreting Referring Expressions in Intergroup Bias
di: Govindarajan, Venkata S, et al.
Pubblicazione: (2024) -
How people talk about each other: Modeling Generalized Intergroup Bias and Emotion
di: Govindarajan, Venkata S, et al.
Pubblicazione: (2022) -
Lil-Bevo: Explorations of Strategies for Training Language Models in More Humanlike Ways
di: Govindarajan, Venkata S, et al.
Pubblicazione: (2023) -
Strategic Dialogue Assessment: The Crooked Path to Innocence
di: Zheng, Anshun Asher, et al.
Pubblicazione: (2025) -
Help! Need Advice on Identifying Advice
di: Govindarajan, Venkata Subrahmanyan, et al.
Pubblicazione: (2020)