Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Rongwu, Zhou, Zi'an, Zhang, Tianwei, Qi, Zehan, Yao, Su, Xu, Ke, Xu, Wei, Qiu, Han |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
DebateQA: Evaluating Question Answering on Debatable Knowledge
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
Long$^2$RAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall
di: Qi, Zehan, et al.
Pubblicazione: (2024)
di: Qi, Zehan, et al.
Pubblicazione: (2024)
Exploring Chinese Humor Generation: A Study on Two-Part Allegorical Sayings
di: Xu, Rongwu
Pubblicazione: (2024)
di: Xu, Rongwu
Pubblicazione: (2024)
A Day in Their Shoes: Using LLM-Based Perspective-Taking Interactive Fiction to Reduce Stigma Toward Dirty Work
di: Yuan, Xiangzhe, et al.
Pubblicazione: (2025)
di: Yuan, Xiangzhe, et al.
Pubblicazione: (2025)
Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency
di: Liu, Yiran, et al.
Pubblicazione: (2024)
di: Liu, Yiran, et al.
Pubblicazione: (2024)
Let's Put Ourselves in Sally's Shoes: Shoes-of-Others Prefilling Improves Theory of Mind in Large Language Models
di: Shinoda, Kazutoshi, et al.
Pubblicazione: (2025)
di: Shinoda, Kazutoshi, et al.
Pubblicazione: (2025)
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
di: Xu, Rongwu, et al.
Pubblicazione: (2023)
di: Xu, Rongwu, et al.
Pubblicazione: (2023)
Knowledge Conflicts for LLMs: A Survey
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
di: Qi, Xuan, et al.
Pubblicazione: (2025)
di: Qi, Xuan, et al.
Pubblicazione: (2025)
When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
di: Wang, Cheng, et al.
Pubblicazione: (2025)
di: Wang, Cheng, et al.
Pubblicazione: (2025)
Course-Correction: Safety Alignment Using Synthetic Preferences
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
AI Awareness
di: Li, Xiaojian, et al.
Pubblicazione: (2025)
di: Li, Xiaojian, et al.
Pubblicazione: (2025)
Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents
di: Xu, Rongwu, et al.
Pubblicazione: (2025)
di: Xu, Rongwu, et al.
Pubblicazione: (2025)
GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language Models
di: Tang, Kunsheng, et al.
Pubblicazione: (2024)
di: Tang, Kunsheng, et al.
Pubblicazione: (2024)
McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
di: Lan, Tian, et al.
Pubblicazione: (2025)
di: Lan, Tian, et al.
Pubblicazione: (2025)
Walk in Their Shoes to Navigate Your Own Path: Learning About Procrastination Through A Serious Game
di: Zhang, Runhua, et al.
Pubblicazione: (2025)
di: Zhang, Runhua, et al.
Pubblicazione: (2025)
DR.GAP: Mitigating Bias in Large Language Models using Gender-Aware Prompting with Demonstration and Reasoning
di: Qiu, Hongye, et al.
Pubblicazione: (2025)
di: Qiu, Hongye, et al.
Pubblicazione: (2025)
When2Speak: A Dataset for Temporal Participation and Turn-Taking in Multi-Party Conversations for Large Language Models
di: Nama, Vihaan, et al.
Pubblicazione: (2026)
di: Nama, Vihaan, et al.
Pubblicazione: (2026)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
Sing it, Narrate it: Quality Musical Lyrics Translation
di: Ye, Zhuorui, et al.
Pubblicazione: (2024)
di: Ye, Zhuorui, et al.
Pubblicazione: (2024)
Detecting LLM-Generated Spam Reviews by Integrating Language Model Embeddings and Graph Neural Network
di: Liu, Xin, et al.
Pubblicazione: (2025)
di: Liu, Xin, et al.
Pubblicazione: (2025)
When and How to Integrate Multimodal Large Language Models in College Psychotherapy: Perspectives from Multi-stakeholders
di: Wang, Jiyao, et al.
Pubblicazione: (2025)
di: Wang, Jiyao, et al.
Pubblicazione: (2025)
Characterization of Political Polarized Users Attacked by Language Toxicity on Twitter
di: Xu, Wentao
Pubblicazione: (2024)
di: Xu, Wentao
Pubblicazione: (2024)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
di: Yang, Shujian, et al.
Pubblicazione: (2025)
di: Yang, Shujian, et al.
Pubblicazione: (2025)
Evaluation of Bias Towards Medical Professionals in Large Language Models
di: Chen, Xi, et al.
Pubblicazione: (2024)
di: Chen, Xi, et al.
Pubblicazione: (2024)
Take its Essence, Discard its Dross! Debiasing for Toxic Language Detection via Counterfactual Causal Effect
di: Lu, Junyu, et al.
Pubblicazione: (2024)
di: Lu, Junyu, et al.
Pubblicazione: (2024)
Gender and Race Bias in Consumer Product Recommendations by Large Language Models
di: Xu, Ke, et al.
Pubblicazione: (2026)
di: Xu, Ke, et al.
Pubblicazione: (2026)
Rehearsing Answers to Probable Questions with Perspective-Taking
di: Shih, Yung-Yu, et al.
Pubblicazione: (2024)
di: Shih, Yung-Yu, et al.
Pubblicazione: (2024)
Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
di: Dutta, Arka, et al.
Pubblicazione: (2023)
di: Dutta, Arka, et al.
Pubblicazione: (2023)
$μ$KE: Matryoshka Unstructured Knowledge Editing of Large Language Models
di: Su, Zian, et al.
Pubblicazione: (2025)
di: Su, Zian, et al.
Pubblicazione: (2025)
A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
di: Xu, Yuemei, et al.
Pubblicazione: (2024)
di: Xu, Yuemei, et al.
Pubblicazione: (2024)
From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models from the Lens of Social Relationship
di: Xu, Yue, et al.
Pubblicazione: (2025)
di: Xu, Yue, et al.
Pubblicazione: (2025)
Take Care of Your Prompt Bias! Investigating and Mitigating Prompt Bias in Factual Knowledge Extraction
di: Xu, Ziyang, et al.
Pubblicazione: (2024)
di: Xu, Ziyang, et al.
Pubblicazione: (2024)
Global Position Aware Group Choreography using Large Language Model
di: Pang, Haozhou, et al.
Pubblicazione: (2025)
di: Pang, Haozhou, et al.
Pubblicazione: (2025)
Expert-Guided Extinction of Toxic Tokens for Debiased Generation
di: Sun, Xueyao, et al.
Pubblicazione: (2024)
di: Sun, Xueyao, et al.
Pubblicazione: (2024)
Promoting Equality in Large Language Models: Identifying and Mitigating the Implicit Bias based on Bayesian Theory
di: Deng, Yongxin, et al.
Pubblicazione: (2024)
di: Deng, Yongxin, et al.
Pubblicazione: (2024)
Communication Bias in Large Language Models: A Regulatory Perspective
di: Kuenzler, Adrian, et al.
Pubblicazione: (2025)
di: Kuenzler, Adrian, et al.
Pubblicazione: (2025)
ShoeModel: Learning to Wear on the User-specified Shoes via Diffusion Model
di: Chen, Binghui, et al.
Pubblicazione: (2024)
di: Chen, Binghui, et al.
Pubblicazione: (2024)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
di: Xu, Xin, et al.
Pubblicazione: (2025)
di: Xu, Xin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
di: Xu, Rongwu, et al.
Pubblicazione: (2024) -
DebateQA: Evaluating Question Answering on Debatable Knowledge
di: Xu, Rongwu, et al.
Pubblicazione: (2024) -
Long$^2$RAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall
di: Qi, Zehan, et al.
Pubblicazione: (2024) -
Exploring Chinese Humor Generation: A Study on Two-Part Allegorical Sayings
di: Xu, Rongwu
Pubblicazione: (2024) -
A Day in Their Shoes: Using LLM-Based Perspective-Taking Interactive Fiction to Reduce Stigma Toward Dirty Work
di: Yuan, Xiangzhe, et al.
Pubblicazione: (2025)