Thinking Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Furniturewala, Shaz, Jandial, Surgan, Java, Abhinav, Banerjee, Pragyan, Shahid, Simra, Bhatia, Sumit, Jaidka, Kokil |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Turn-Level Empathy Prediction Using Psychological Indicators
by: Furniturewala, Shaz, et al.
Published: (2024)
by: Furniturewala, Shaz, et al.
Published: (2024)
Learning Through Dialogue: Engagement and Efficacy Matter More Than Explanations
by: Furniturewala, Shaz, et al.
Published: (2026)
by: Furniturewala, Shaz, et al.
Published: (2026)
Impact of Decoding Methods on Human Alignment of Conversational LLMs
by: Furniturewala, Shaz, et al.
Published: (2024)
by: Furniturewala, Shaz, et al.
Published: (2024)
LEAST: "Local" text-conditioned image style transfer
by: Singh, Silky, et al.
Published: (2024)
by: Singh, Silky, et al.
Published: (2024)
Beyond Text: Leveraging Multi-Task Learning and Cognitive Appraisal Theory for Post-Purchase Intention Analysis
by: Yeo, Gerard Christopher, et al.
Published: (2024)
by: Yeo, Gerard Christopher, et al.
Published: (2024)
Towards Operationalizing Right to Data Protection
by: Java, Abhinav, et al.
Published: (2024)
by: Java, Abhinav, et al.
Published: (2024)
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content
by: Furniturewala, Shaz, et al.
Published: (2025)
by: Furniturewala, Shaz, et al.
Published: (2025)
Beyond Context to Cognitive Appraisal: Emotion Reasoning as a Theory of Mind Benchmark for Large Language Models
by: Yeo, Gerard Christopher, et al.
Published: (2025)
by: Yeo, Gerard Christopher, et al.
Published: (2025)
The MediaSpin Dataset: Post-Publication News Headline Edits Annotated for Media Bias
by: Verma, Preetika, et al.
Published: (2024)
by: Verma, Preetika, et al.
Published: (2024)
Reading Between the Lines: How Electronic Nonverbal Cues shape Emotion Decoding
by: Kumar, Taara, et al.
Published: (2026)
by: Kumar, Taara, et al.
Published: (2026)
Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation
by: Churina, Svetlana, et al.
Published: (2024)
by: Churina, Svetlana, et al.
Published: (2024)
Do You Trust Me? Cognitive-Affective Signatures of Trustworthiness in Large Language Models
by: Yeo, Gerard, et al.
Published: (2025)
by: Yeo, Gerard, et al.
Published: (2025)
On the Limitations of Steering in Language Model Alignment
by: Niranjan, Chebrolu, et al.
Published: (2025)
by: Niranjan, Chebrolu, et al.
Published: (2025)
"Reasoning" with Rhetoric: On the Style-Evidence Tradeoff in LLM-Generated Counter-Arguments
by: Verma, Preetika, et al.
Published: (2024)
by: Verma, Preetika, et al.
Published: (2024)
ReEdit: Multimodal Exemplar-Based Image Editing with Diffusion Models
by: Srivastava, Ashutosh, et al.
Published: (2024)
by: Srivastava, Ashutosh, et al.
Published: (2024)
S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models
by: Shukla, Nitish, et al.
Published: (2026)
by: Shukla, Nitish, et al.
Published: (2026)
Conversations: Love Them, Hate Them, Steer Them
by: Chebrolu, Niranjan, et al.
Published: (2025)
by: Chebrolu, Niranjan, et al.
Published: (2025)
GitSearch: Enhancing Community Notes Generation with Gap-Informed Targeted Search
by: Singh, Sahajpreet, et al.
Published: (2026)
by: Singh, Sahajpreet, et al.
Published: (2026)
Towards Efficient Exemplar Based Image Editing with Multimodal VLMs
by: Jadhav, Avadhoot, et al.
Published: (2025)
by: Jadhav, Avadhoot, et al.
Published: (2025)
From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation
by: Chebrolu, Niranjan, et al.
Published: (2025)
by: Chebrolu, Niranjan, et al.
Published: (2025)
LLMs and Finetuning: Benchmarking cross-domain performance for hate speech detection
by: Nasir, Ahmad, et al.
Published: (2023)
by: Nasir, Ahmad, et al.
Published: (2023)
Labels or Input? Rethinking Augmentation in Multimodal Hate Detection
by: Singh, Sahajpreet, et al.
Published: (2025)
by: Singh, Sahajpreet, et al.
Published: (2025)
Predicting Sentence-Level Factuality of News and Bias of Media Outlets
by: Vargas, Francielle, et al.
Published: (2023)
by: Vargas, Francielle, et al.
Published: (2023)
Disentangling Codemixing in Chats: The NUS ABC Codemixed Corpus
by: Churina, Svetlana, et al.
Published: (2025)
by: Churina, Svetlana, et al.
Published: (2025)
Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
by: Ouyang, Rongxin, et al.
Published: (2024)
by: Ouyang, Rongxin, et al.
Published: (2024)
CommunityFact: A Dynamic, Multilingual, Multi-domain Benchmark for Misinformation Detection in the Wild
by: Singh, Sahajpreet, et al.
Published: (2026)
by: Singh, Sahajpreet, et al.
Published: (2026)
PHAnToM: Persona-based Prompting Has An Effect on Theory-of-Mind Reasoning in Large Language Models
by: Tan, Fiona Anting, et al.
Published: (2024)
by: Tan, Fiona Anting, et al.
Published: (2024)
FrugalRAG: Less is More in RL Finetuning for Multi-Hop Question Answering
by: Java, Abhinav, et al.
Published: (2025)
by: Java, Abhinav, et al.
Published: (2025)
Unlocking Structured Thinking in Language Models with Cognitive Prompting
by: Kramer, Oliver, et al.
Published: (2024)
by: Kramer, Oliver, et al.
Published: (2024)
Rethinking Prompt-based Debiasing in Large Language Models
by: Yang, Xinyi, et al.
Published: (2025)
by: Yang, Xinyi, et al.
Published: (2025)
POSIX: A Prompt Sensitivity Index For Large Language Models
by: Chatterjee, Anwoy, et al.
Published: (2024)
by: Chatterjee, Anwoy, et al.
Published: (2024)
On Bias and Fairness in NLP: Investigating the Impact of Bias and Debiasing in Language Models on the Fairness of Toxicity Detection
by: Elsafoury, Fatma, et al.
Published: (2023)
by: Elsafoury, Fatma, et al.
Published: (2023)
Causal Prompting: Debiasing Large Language Model Prompting based on Front-Door Adjustment
by: Zhang, Congzhi, et al.
Published: (2024)
by: Zhang, Congzhi, et al.
Published: (2024)
Althea: Human-AI Collaboration for Fact-Checking and Critical Reasoning
by: Churina, Svetlana, et al.
Published: (2025)
by: Churina, Svetlana, et al.
Published: (2025)
Heterogeneity in Formal Linguistic Competence of Language Models: Is Data the Real Bottleneck?
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2026)
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2026)
Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow
by: Du, Y., et al.
Published: (2025)
by: Du, Y., et al.
Published: (2025)
Fast-Slow-Thinking: Complex Task Solving with Large Language Models
by: Sun, Yiliu, et al.
Published: (2025)
by: Sun, Yiliu, et al.
Published: (2025)
Enhancing Creativity in Large Language Models through Associative Thinking Strategies
by: Mehrotra, Pronita, et al.
Published: (2024)
by: Mehrotra, Pronita, et al.
Published: (2024)
Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector
by: Yang, Haoyan, et al.
Published: (2025)
by: Yang, Haoyan, et al.
Published: (2025)
Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and Translation
by: Limisiewicz, Tomasz, et al.
Published: (2025)
by: Limisiewicz, Tomasz, et al.
Published: (2025)
Similar Items
-
Turn-Level Empathy Prediction Using Psychological Indicators
by: Furniturewala, Shaz, et al.
Published: (2024) -
Learning Through Dialogue: Engagement and Efficacy Matter More Than Explanations
by: Furniturewala, Shaz, et al.
Published: (2026) -
Impact of Decoding Methods on Human Alignment of Conversational LLMs
by: Furniturewala, Shaz, et al.
Published: (2024) -
LEAST: "Local" text-conditioned image style transfer
by: Singh, Silky, et al.
Published: (2024) -
Beyond Text: Leveraging Multi-Task Learning and Cognitive Appraisal Theory for Post-Purchase Intention Analysis
by: Yeo, Gerard Christopher, et al.
Published: (2024)