Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yuchen, Sun, Edward, Zhu, Kaijie, Lian, Jianxun, Hernandez-Orallo, Jose, Caliskan, Aylin, Wang, Jindong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Taxonomy of Stereotype Content in Large Language Models
by: Nicolas, Gandalf, et al.
Published: (2024)
by: Nicolas, Gandalf, et al.
Published: (2024)
ChatGPT Perpetuates Gender Bias in Machine Translation and Ignores Non-Gendered Pronouns: Findings across Bengali and Five other Low-Resource Languages
by: Ghosh, Sourojit, et al.
Published: (2023)
by: Ghosh, Sourojit, et al.
Published: (2023)
Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval
by: Wilson, Kyra, et al.
Published: (2024)
by: Wilson, Kyra, et al.
Published: (2024)
"I don't see myself represented here at all": User Experiences of Stable Diffusion Outputs Containing Representational Harms across Gender Identities and Nationalities
by: Ghosh, Sourojit, et al.
Published: (2024)
by: Ghosh, Sourojit, et al.
Published: (2024)
Bias Amplification in Stable Diffusion's Representation of Stigma Through Skin Tones and Their Homogeneity
by: Wilson, Kyra, et al.
Published: (2025)
by: Wilson, Kyra, et al.
Published: (2025)
Beyond One-Way Influence: Bidirectional Opinion Dynamics in Multi-Turn Human-LLM Interactions
by: Jiang, Yuyang, et al.
Published: (2025)
by: Jiang, Yuyang, et al.
Published: (2025)
Do Generative AI Models Output Harm while Representing Non-Western Cultures: Evidence from A Community-Centered Approach
by: Ghosh, Sourojit, et al.
Published: (2024)
by: Ghosh, Sourojit, et al.
Published: (2024)
'Person' == Light-skinned, Western Man, and Sexualization of Women of Color: Stereotypes in Stable Diffusion
by: Ghosh, Sourojit, et al.
Published: (2023)
by: Ghosh, Sourojit, et al.
Published: (2023)
VEAT Quantifies Implicit Associations in Text-to-Video Generator Sora and Reveals Challenges in Bias Mitigation
by: Sun, Yongxu, et al.
Published: (2026)
by: Sun, Yongxu, et al.
Published: (2026)
No Thoughts Just AI: Biased LLM Hiring Recommendations Alter Human Decision Making and Limit Human Autonomy
by: Wilson, Kyra, et al.
Published: (2025)
by: Wilson, Kyra, et al.
Published: (2025)
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
by: Voudouris, Konstantinos, et al.
Published: (2026)
by: Voudouris, Konstantinos, et al.
Published: (2026)
Evaluating General-Purpose AI with Psychometrics
by: Wang, Xiting, et al.
Published: (2023)
by: Wang, Xiting, et al.
Published: (2023)
Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation
by: Gueorguieva, Anna-Maria, et al.
Published: (2025)
by: Gueorguieva, Anna-Maria, et al.
Published: (2025)
Designing AI-Agents with Personalities: A Psychometric Approach
by: Huang, Muhua, et al.
Published: (2024)
by: Huang, Muhua, et al.
Published: (2024)
Computational Multi-Agents Society Experiments: Social Modeling Framework Based on Generative Agents
by: Zhang, Hanzhong, et al.
Published: (2025)
by: Zhang, Hanzhong, et al.
Published: (2025)
How Should AI Safety Benchmarks Benchmark Safety?
by: Yu, Cheng, et al.
Published: (2026)
by: Yu, Cheng, et al.
Published: (2026)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
by: Li, Miles Q., et al.
Published: (2026)
by: Li, Miles Q., et al.
Published: (2026)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
by: Gomaa, Amr, et al.
Published: (2025)
by: Gomaa, Amr, et al.
Published: (2025)
Evaluating LLM Safety Across Child Development Stages: A Simulated Agent Approach
by: Murali, Abhejay, et al.
Published: (2025)
by: Murali, Abhejay, et al.
Published: (2025)
A Multi-Agent Approach to Validate and Refine LLM-Generated Personalized Math Problems
by: Ikram, Fareya, et al.
Published: (2026)
by: Ikram, Fareya, et al.
Published: (2026)
A Multimodal Manufacturing Safety Chatbot: Knowledge Base Design, Benchmark Development, and Evaluation of Multiple RAG Approaches
by: Singh, Ryan, et al.
Published: (2025)
by: Singh, Ryan, et al.
Published: (2025)
ChatGPT as Research Scientist: Probing GPT's Capabilities as a Research Librarian, Research Ethicist, Data Generator and Data Predictor
by: Lehr, Steven A., et al.
Published: (2024)
by: Lehr, Steven A., et al.
Published: (2024)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
by: Li, Chance Jiajie, et al.
Published: (2025)
by: Li, Chance Jiajie, et al.
Published: (2025)
DM-Bench: Benchmarking LLMs for Personalized Decision Making in Diabetes Management
by: Cardei, Maria Ana, et al.
Published: (2025)
by: Cardei, Maria Ana, et al.
Published: (2025)
Refined floor diagrams relative to a conic and Caporaso–Harris‐type formula
by: Yanqiao Ding, et al.
Published: (2026)
by: Yanqiao Ding, et al.
Published: (2026)
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
The Impact of Parental Divorce on Children Through the Mothers' Lens: A Phenomenological Study
by: Zeynep Kisecik Sengul, et al.
Published: (2026)
by: Zeynep Kisecik Sengul, et al.
Published: (2026)
Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
by: Fitz, Stephen, et al.
Published: (2025)
by: Fitz, Stephen, et al.
Published: (2025)
A Survey on Responsible Generative AI: What to Generate and What Not
by: Gu, Jindong
Published: (2024)
by: Gu, Jindong
Published: (2024)
Safety Cases: A Scalable Approach to Frontier AI Safety
by: Hilton, Benjamin, et al.
Published: (2025)
by: Hilton, Benjamin, et al.
Published: (2025)
AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
by: Zeng, Yi, et al.
Published: (2024)
by: Zeng, Yi, et al.
Published: (2024)
WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making
by: Li, Zongjie, et al.
Published: (2026)
by: Li, Zongjie, et al.
Published: (2026)
EpiPlanAgent: Agentic Automated Epidemic Response Planning
by: Mao, Kangkun, et al.
Published: (2025)
by: Mao, Kangkun, et al.
Published: (2025)
The Better Angels of Machine Personality: How Personality Relates to LLM Safety
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Auditing Agent Harness Safety
by: Liu, Chengzhi, et al.
Published: (2026)
by: Liu, Chengzhi, et al.
Published: (2026)
AI Safety Frameworks Should Include Procedures for Model Access Decisions
by: Kembery, Edward, et al.
Published: (2024)
by: Kembery, Edward, et al.
Published: (2024)
Benchmarking LLMs for Community Governance Simulation with Life-history Narratives
by: Chen, Xu, et al.
Published: (2026)
by: Chen, Xu, et al.
Published: (2026)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
by: Jin, Yiqiao, et al.
Published: (2024)
by: Jin, Yiqiao, et al.
Published: (2024)
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
by: Potham, Ram
Published: (2025)
by: Potham, Ram
Published: (2025)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
by: Shi, Yuzhen, et al.
Published: (2026)
by: Shi, Yuzhen, et al.
Published: (2026)
Similar Items
-
A Taxonomy of Stereotype Content in Large Language Models
by: Nicolas, Gandalf, et al.
Published: (2024) -
ChatGPT Perpetuates Gender Bias in Machine Translation and Ignores Non-Gendered Pronouns: Findings across Bengali and Five other Low-Resource Languages
by: Ghosh, Sourojit, et al.
Published: (2023) -
Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval
by: Wilson, Kyra, et al.
Published: (2024) -
"I don't see myself represented here at all": User Experiences of Stable Diffusion Outputs Containing Representational Harms across Gender Identities and Nationalities
by: Ghosh, Sourojit, et al.
Published: (2024) -
Bias Amplification in Stable Diffusion's Representation of Stigma Through Skin Tones and Their Homogeneity
by: Wilson, Kyra, et al.
Published: (2025)