Evaluating Chinese Large Language Models: The Influence of Persona Assignment on Stereotypes and Safeguards
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Geng, Feng, Li, Bono, Carlo Alberto, Yang, Songbo, Zhu, Mengxiao, Pierri, Francesco |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Comparing diversity, negativity, and stereotypes in Chinese-language AI technologies: an investigation of Baidu, Ernie and Qwen
di: Liu, Geng, et al.
Pubblicazione: (2024)
di: Liu, Geng, et al.
Pubblicazione: (2024)
Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers
di: Liu, Geng, et al.
Pubblicazione: (2025)
di: Liu, Geng, et al.
Pubblicazione: (2025)
Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts
di: Nakanishi, Akito, et al.
Pubblicazione: (2025)
di: Nakanishi, Akito, et al.
Pubblicazione: (2025)
Evaluating open-source Large Language Models for automated fact-checking
di: Fontana, Nicolo', et al.
Pubblicazione: (2025)
di: Fontana, Nicolo', et al.
Pubblicazione: (2025)
Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups
di: Liu, Geng, et al.
Pubblicazione: (2025)
di: Liu, Geng, et al.
Pubblicazione: (2025)
Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Mindset in Large Language Models
di: Corso, Francesco, et al.
Pubblicazione: (2025)
di: Corso, Francesco, et al.
Pubblicazione: (2025)
Static and Dynamic Strategies for Influencing Opinions in Social Networks
di: Tarantino, Paolo, et al.
Pubblicazione: (2026)
di: Tarantino, Paolo, et al.
Pubblicazione: (2026)
Nicer Than Humans: How do Large Language Models Behave in the Prisoner's Dilemma?
di: Fontana, Nicoló, et al.
Pubblicazione: (2024)
di: Fontana, Nicoló, et al.
Pubblicazione: (2024)
A comparison of online search engine autocompletion in Google and Baidu
di: Liu, Geng, et al.
Pubblicazione: (2024)
di: Liu, Geng, et al.
Pubblicazione: (2024)
Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency
di: Liu, Yiran, et al.
Pubblicazione: (2024)
di: Liu, Yiran, et al.
Pubblicazione: (2024)
Evaluating AI capabilities in detecting conspiracy theories on YouTube
di: La Rocca, Leonardo, et al.
Pubblicazione: (2025)
di: La Rocca, Leonardo, et al.
Pubblicazione: (2025)
$\texttt{ModSCAN}$: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities
di: Jiang, Yukun, et al.
Pubblicazione: (2024)
di: Jiang, Yukun, et al.
Pubblicazione: (2024)
Towards an Automated Framework to Audit Youth Safety on TikTok
di: Xue, Linda, et al.
Pubblicazione: (2025)
di: Xue, Linda, et al.
Pubblicazione: (2025)
From Content to Audience: A Multimodal Annotation Framework for Broadcast Television Analytics
di: Cupini, Paolo, et al.
Pubblicazione: (2026)
di: Cupini, Paolo, et al.
Pubblicazione: (2026)
Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models
di: Siddique, Zara, et al.
Pubblicazione: (2024)
di: Siddique, Zara, et al.
Pubblicazione: (2024)
A Taxonomy of Stereotype Content in Large Language Models
di: Nicolas, Gandalf, et al.
Pubblicazione: (2024)
di: Nicolas, Gandalf, et al.
Pubblicazione: (2024)
A Longitudinal Study of Italian and French Reddit Conversations Around the Russian Invasion of Ukraine
di: Corso, Francesco, et al.
Pubblicazione: (2024)
di: Corso, Francesco, et al.
Pubblicazione: (2024)
Addressing Stereotypes in Large Language Models: A Critical Examination and Mitigation
di: Kazi, Fatima
Pubblicazione: (2025)
di: Kazi, Fatima
Pubblicazione: (2025)
Among Us: Language of Conspiracy Theorists on Mainstream Reddit
di: Corso, Francesco, et al.
Pubblicazione: (2025)
di: Corso, Francesco, et al.
Pubblicazione: (2025)
Evaluating Proactive Risk Awareness of Large Language Models
di: Luo, Xuan, et al.
Pubblicazione: (2026)
di: Luo, Xuan, et al.
Pubblicazione: (2026)
Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2024)
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2024)
DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis
di: Vijayaraghavan, Prashanth, et al.
Pubblicazione: (2025)
di: Vijayaraghavan, Prashanth, et al.
Pubblicazione: (2025)
Effects of Algorithmic Visibility on Conspiracy Communities: Reddit after Epstein's 'Suicide'
di: Attanasio, Asja, et al.
Pubblicazione: (2025)
di: Attanasio, Asja, et al.
Pubblicazione: (2025)
SESGO: Spanish Evaluation of Stereotypical Generative Outputs
di: Robles, Melissa, et al.
Pubblicazione: (2025)
di: Robles, Melissa, et al.
Pubblicazione: (2025)
An Empirical Investigation of Gender Stereotype Representation in Large Language Models: The Italian Case
di: Giachino, Gioele, et al.
Pubblicazione: (2025)
di: Giachino, Gioele, et al.
Pubblicazione: (2025)
Conspiracy theories and where to find them on TikTok
di: Corso, Francesco, et al.
Pubblicazione: (2024)
di: Corso, Francesco, et al.
Pubblicazione: (2024)
What we can learn from TikTok through its Research API
di: Corso, Francesco, et al.
Pubblicazione: (2024)
di: Corso, Francesco, et al.
Pubblicazione: (2024)
Self-Debiasing Large Language Models: Zero-Shot Recognition and Reduction of Stereotypes
di: Gallegos, Isabel O., et al.
Pubblicazione: (2024)
di: Gallegos, Isabel O., et al.
Pubblicazione: (2024)
Student Perspectives on Using a Large Language Model (LLM) for an Assignment on Professional Ethics
di: Grande, Virginia, et al.
Pubblicazione: (2024)
di: Grande, Virginia, et al.
Pubblicazione: (2024)
Multilinguality at the Edge: Developing Language Models for the Global South
di: Miranda, Lester James V., et al.
Pubblicazione: (2026)
di: Miranda, Lester James V., et al.
Pubblicazione: (2026)
SOTOPIA-$Ω$: Dynamic Strategy Injection Learning and Social Instruction Following Evaluation for Social Agents
di: Zhang, Wenyuan, et al.
Pubblicazione: (2025)
di: Zhang, Wenyuan, et al.
Pubblicazione: (2025)
Safeguarding Efficacy in Large Language Models: Evaluating Resistance to Human-Written and Algorithmic Adversarial Prompts
di: Downey-Webb, Tiarnaigh, et al.
Pubblicazione: (2025)
di: Downey-Webb, Tiarnaigh, et al.
Pubblicazione: (2025)
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
di: Greco, Candida M., et al.
Pubblicazione: (2026)
di: Greco, Candida M., et al.
Pubblicazione: (2026)
A Chinese Dataset for Evaluating the Safeguards in Large Language Models
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
A Survey on Stereotype Detection in Natural Language Processing
di: Cignarella, Alessandra Teresa, et al.
Pubblicazione: (2025)
di: Cignarella, Alessandra Teresa, et al.
Pubblicazione: (2025)
A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI Evaluations
di: Davani, Aida, et al.
Pubblicazione: (2025)
di: Davani, Aida, et al.
Pubblicazione: (2025)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
di: Kiet, Huynh Trung, et al.
Pubblicazione: (2026)
di: Kiet, Huynh Trung, et al.
Pubblicazione: (2026)
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
di: Costa, Davi Bastos, et al.
Pubblicazione: (2025)
di: Costa, Davi Bastos, et al.
Pubblicazione: (2025)
Influence of Personality Traits on Plagiarism Through Collusion in Programming Assignments
di: PD, Parthasarathy, et al.
Pubblicazione: (2024)
di: PD, Parthasarathy, et al.
Pubblicazione: (2024)
Evaluation of Bias Towards Medical Professionals in Large Language Models
di: Chen, Xi, et al.
Pubblicazione: (2024)
di: Chen, Xi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Comparing diversity, negativity, and stereotypes in Chinese-language AI technologies: an investigation of Baidu, Ernie and Qwen
di: Liu, Geng, et al.
Pubblicazione: (2024) -
Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers
di: Liu, Geng, et al.
Pubblicazione: (2025) -
Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts
di: Nakanishi, Akito, et al.
Pubblicazione: (2025) -
Evaluating open-source Large Language Models for automated fact-checking
di: Fontana, Nicolo', et al.
Pubblicazione: (2025) -
Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups
di: Liu, Geng, et al.
Pubblicazione: (2025)